Multi-dimensional pixel fusion method and system based on multi-source perception
Through multi-source perception technology and multi-dimensional pixel fusion method, combined with user voice and facial expressions, the problem of low interaction accuracy of existing humanoid robots is solved, and more accurate behavioral operation instructions and better user experience are achieved.
Patent Information
- Application Number
- CN202510423254.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-07
AI Technical Summary
When interacting with users, existing humanoid robots rely on a single interaction method and cannot provide multi-directional feedback based on the user's facial expressions and tone, resulting in low interaction accuracy and unable to satisfy the user's interactive experience.
The multi-dimensional pixel fusion method based on multi-source perception is adopted to collect data through sensors such as visible light cameras, infrared cameras and lidars, and combine tone feedback during user voice, multi-dimensional pixel information fusion processing is carried out to generate more accurate behavioral operation instructions.
It improves interaction efficiency, satisfies the user's interactive experience, makes behavioral operation instructions more accurate, and enhances the interactive experience between the robot and the user.
Smart Images

Figure CN119942291A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of pixel fusion technology, and in particular to a multi-dimensional pixel fusion method and system based on multi-source perception. Background Art
[0002] With the increasing promotion of humanoid robots and the increasing maturity of technology, more and more users are attracted to actively chat and interact with humanoid robots. Common humanoid robots provide interactive feedback by collecting user interaction information, and generally rely on a single interactive method during the collection process. For example, simply outputting text or voice feedback to the user does not provide multi-faceted feedback based on the user's facial expressions or tone, resulting in low interaction accuracy, that is, the user's interactive experience cannot be satisfied due to low interaction efficiency.
[0003] Therefore, the present invention proposes a multi-dimensional pixel fusion method and system based on multi-source perception. Summary of the invention
[0004] The present invention provides a multi-dimensional pixel fusion method and system based on multi-source perception, which are used to determine multi-dimensional pixel information from multi-source perception, and combine the influence of voice feedback in the user's voice process to ensure the reliability of multi-dimensional pixel fusion, thereby making behavioral operation instructions more accurate, thereby improving interaction efficiency and satisfying the user's interactive experience.
[0005] The present invention provides a multi-dimensional pixel fusion method based on multi-source perception, comprising: Step 1: Control the sensors set on the robot to collect data, wherein the sensors include: a visible light camera, an infrared camera and a laser radar; Step 2: Unify the coordinates and align the time sequence of the collected data of each sensor to obtain the collection array of each collection point; Step 3: Arrange and process the collected data of each type of sensor in spatial order to obtain a corresponding collected image, and perform feature extraction on the collected image according to a feature extraction method matching the corresponding type of sensor; Step 4: Align the extracted features under each type of sensor, and perform multi-dimensional pixel fusion in combination with the acquisition array of the acquisition point and the tone attribute of the target user to obtain behavioral operation instructions and output them to the control unit of the robot to perform corresponding operations.
[0006] Preferably, before the sensors provided on the control robot itself are controlled to collect data, the following steps are included: Capturing a trigger instruction of the robot from external information, and determining an instruction type of the trigger instruction; A start unit consistent with the instruction type is matched from a type-start database, and the robot is started in a power-on state based on the start unit.
[0007] Preferably, controlling the sensors provided on the robot itself to collect data includes: When the robot is in a powered-on state, the robot collects an interactive input set of a target user currently received by the robot, and interactively analyzes the interactive input set to set a collection cycle for different types of sensors; The sensors of the corresponding type are controlled to collect data according to the collection period to obtain an initial set of each sensor.
[0008] Preferably, setting a collection period for different types of sensors includes: Determining an input type of the interactive input set, and when the input type is voice input, respectively obtaining a response curve of each microphone collecting the user's voice; In combination with the setting position of the microphone of the robot itself and the current distance of the sound source of the target user, the response curves are sequentially sorted and aligned to obtain the response difference at each time point; Based on the response difference and time delay, a compensation adjustment is performed on the second curve which is the second closest, and waveform superposition processing is performed on the first curve which is the closest; Performing speech noise reduction processing on the superimposed speech to obtain the converted text; Extracting keywords from the converted text; Determine the duration of the process of the robot receiving the user's voice and performing semantic analysis on the user's voice, and at the same time, determine the voice duration of the user's voice; Obtaining a tone complexity coefficient according to the speech duration, the process duration, the tone information of the speech after the noise reduction process, and the number of words in the converted text; Retrieving a tone feature model that matches the tone complexity coefficient from a coefficient-feature comparison table to analyze the speech after noise reduction processing, and determining a tone annotation for each keyword; Merging the tone annotation with the extracted keywords to obtain a first annotation; When the input type is text input, extract keywords from the input text to obtain a second annotation; The interaction intention of the target user is determined according to the annotation, and the collection cycles of different types of sensors are matched from the intention comparison table.
[0009] Preferably, the extracted features under each type of sensor are aligned, and multi-dimensional pixel fusion is performed in combination with the acquisition array of the acquisition points and the tone attribute of the target user, including: Constructing a gradient vector for each pixel point in the acquisition array at each acquisition point, wherein the gradient vector includes a brightness gradient component of the corresponding pixel point under different types of sensors; Determine the information saturation coefficient of each pixel according to the gradient vector; When the information saturation coefficient is greater than a preset coefficient, the corresponding collection point is retained, otherwise, the corresponding collection point is removed; The extracted features of each type of sensor are converted into the spatial domain to obtain a spatial array of each spatial point. When there are remaining points that are consistent with the corresponding spatial points, the spatial array and the acquisition array under the consistent points are pixel-fused according to the tone attribute of the target user to obtain a fused array and retain it; When there are no corresponding spatial points that are consistent with the remaining points, the acquisition array corresponding to the remaining points is retained; Analyze the reserved array of each point to obtain behavioral operation instructions.
[0010] Preferably, pixel fusion is performed on the spatial array and the acquisition array at the consistent point according to the tone attribute of the target user to obtain a fused array, including: Based on the spatial array and the acquisition array, construct an initial matrix; When the tone attribute is that the input type of the target user's interaction with the robot is unrelated to speech, an average value of each column element in the initial matrix is calculated to obtain a first array; When the tone attribute is an input type related to speech for the target user to interact with the robot, determining the facial area of the target user on which the consistent point acts according to the comparison relationship between the consistent point and the face of the target user, and setting simulation parameters for the consistent point according to the expression control database of the facial area; Extract the current tone of the target user at each time point during the interaction with the robot, and construct facial change information for each consistent point in combination with the user tone portrait of the target user and simulation parameters, and then construct a facial array based on the tone; The face array is appended to the first row of the initial matrix, and combined with the weight assigned to each row vector in the initial matrix, to obtain a second array; Among them, the first array and the second array are corresponding fused arrays.
[0011] Preferably, the reserved array of each point is analyzed to obtain the behavior operation instruction, including: Divide the retained array of each point based on the block division criteria of the user's face to obtain the current representation of each block; Based on all current representations, the block weight of each divided block and the user intention determined by the user voice of the target user, a behavior operation instruction is obtained from the intention-combination-instruction comparison table.
[0012] The present invention provides a multi-dimensional pixel fusion system based on multi-source perception, comprising: A data acquisition module, used to control various sensors set by the robot itself to collect data, wherein the sensors include: a visible light camera, an infrared camera and a laser radar; The alignment processing module is used to unify the coordinates and align the timing of the collected data of each sensor to obtain the collection array of each collection point; A feature extraction module is used to arrange and process the collected data of each type of sensor in spatial order to obtain a corresponding collected image, and perform feature extraction on the collected image according to a feature extraction method matching the corresponding type of sensor; The multi-dimensional pixel fusion module is used to align the extracted features under each type of sensor, and perform multi-dimensional pixel fusion in combination with the acquisition array of the acquisition points and the tone attributes of the target user, obtain behavioral operation instructions and output them to the control unit of the robot to perform corresponding operations.
[0013] Compared with the prior art, the present invention has the following beneficial effects: Starting from multi-source perception, we determine the multi-dimensional pixel information, and combine it with the influence of the user's voice tone feedback during the voice process to ensure the reliability of multi-dimensional pixel fusion, thereby making the behavioral operation instructions more accurate, improving the interaction efficiency and satisfying the user's interactive experience.
[0014] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.
[0015] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 This is a flow chart of a multi-dimensional pixel fusion method based on multi-source perception in an embodiment of the present invention; Figure 2 This is a structural diagram of a multi-dimensional pixel fusion system based on multi-source perception in an embodiment of the present invention; Figure 34 is an external view of a robot in an embodiment of the present invention. DETAILED DESCRIPTION
[0017] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0018] The present invention provides a multi-dimensional pixel fusion method based on multi-source perception, such as Figure 1 As shown, including: Step 1: Control the sensors set on the robot to collect data, wherein the sensors include: a visible light camera, an infrared camera and a laser radar; Step 2: Unify the coordinates and align the time sequence of the collected data of each sensor to obtain the collection array of each collection point; Step 3: Arrange and process the collected data of each type of sensor in spatial order to obtain a corresponding collected image, and perform feature extraction on the collected image according to a feature extraction method matching the corresponding type of sensor; Step 4: Align the extracted features under each type of sensor, and perform multi-dimensional pixel fusion in combination with the acquisition array of the acquisition point and the tone attribute of the target user to obtain behavioral operation instructions and output them to the control unit of the robot to perform corresponding operations.
[0019] In this embodiment, the robot is an existing robot that can mainly interact with humans, and the robot is designed to be a female image, 168 cm tall, with black hair and black eyes. Two types of robot are available: a walking operation model and an optional chassis operation model. The appearance incorporates Chinese elements to bring a unique experience to users.
[0020] The appearance of the robots is integrated with Chinese elements, such as blue and white porcelain, auspicious clouds, and Shijingshan. The cheongsam model only needs to design the appearance of the chassis, and the head refers to the image proportion of the original IP of Huanhuan, without redesigning the appearance. Figure 3 shown.
[0021] In this embodiment, the functions of the robot are as follows, including: Facial expression function: The robot can randomly blink, move its eyeballs, and make other expressions. Using the remote control, it can also be controlled to smile, yawn, frown, and make other rich expressions, vividly showing emotional changes.
[0022] Upper body movement function: Through the remote control, you can wave to say hello, shake hands, swing your arms randomly, turn your head and other movements to meet the needs of various social interaction scenarios.
[0023] Voice chat function: A voice dialogue system based on a large model that supports setting a specific persona. The robot will chat and answer questions in the tone of the set persona, with specific voice wake-up and interruption functions, and can also play pre-recorded audio.
[0024] Visual perception function: With the help of the camera built into the eyeball, the robot can observe the environment and chat with people based on it, but the conversation delay is about 3-5 seconds.
[0025] Bipedal walking function: The remote control can be used to conveniently control the robot's legs to move forward, backward, and turn to achieve flexible walking. Both the bipedal walking of the walking operation model and the chassis walking of the chassis operation model have good controllability. The terrain adaptability of the chassis operation model further ensures stable movement in complex scenes.
[0026] In this embodiment, the behavioral operation instructions are facial emotion interaction instructions, voice chat interaction instructions, body movement interaction instructions, walking control instructions, etc. between the robot and the user.
[0027] In this embodiment, multi-source sensing multi-dimensional pixel fusion refers to a technology that fuses pixel data containing information of different dimensions obtained from multiple different sensors (such as cameras, radars, etc.) to generate a richer, more accurate and comprehensive image or data representation. For example, the pixel information of a visible light image is fused with the pixel information of an infrared image so that the fused image contains both details under visible light and characteristic information such as temperature under infrared light.
[0028] In this embodiment, coordinate unification and time alignment are performed to ensure that pixels of different source images accurately correspond in space. For example, corresponding points of the same physical position in different images are found through feature point matching and other methods, and the images are transformed by translation, rotation, scaling and other transformations to align them in space.
[0029] In this embodiment, representative features are extracted from pixels of each source image, such as extracting features such as edges and textures from visible light images, and extracting geometric features such as distances to objects from depth images.
[0030] In this embodiment, the pixel fusion algorithm includes weighted average fusion, wavelet transform-based fusion, sparse representation-based fusion, etc. For example, in weighted average fusion, a weight is assigned to each pixel according to factors such as the reliability of different source images, and then the weighted sum is performed to obtain the fused pixel value.
[0031] In this embodiment, the collection point is the point where the face of the target user is collected, and the collected data is pixel data related to the user's face.
[0032] In this embodiment, the acquisition array = {pixel information acquired by sensors of each type at corresponding acquisition points}.
[0033] In this embodiment, because there is at least one sensor of each type and the sensors are installed at different positions, a complete captured image can be obtained in the process of arranging in spatial order, that is, each type of sensor corresponds to a captured image at the corresponding moment, which can be a facial image of the user, and the target user refers to the user who interacts with the robot.
[0034] In this embodiment, PyTorch can be used to extract features of visible light images, local invariant features can be used to extract features of infrared images, and the DBSCAN algorithm can be used to extract features of lidar images.
[0035] In this embodiment, feature alignment processing refers to position alignment processing of the features extracted at each time point. This is because only valuable points are retained during the feature extraction process. Therefore, feature information of related spatial points can be obtained after alignment processing.
[0036] In this embodiment, the tone attribute of the target user includes: whether there is voice interaction between the user and the robot, or whether there is no voice interaction between the user and the robot.
[0037] In this embodiment, multi-dimensional pixel fusion refers to fusing pixels existing at corresponding positions to facilitate subsequent analysis and obtain behavioral operation instructions.
[0038] The beneficial effect of the above technical solution is: starting from multi-source perception to determine multi-dimensional pixel information, and combining the influence of tone feedback in the user's voice process to ensure the reliability of multi-dimensional pixel fusion, thereby making behavioral operation instructions more accurate, to improve interaction efficiency and satisfy the user's interactive experience.
[0039] The present invention provides a multi-dimensional pixel fusion method based on multi-source perception, which includes: Capturing a trigger instruction of the robot from external information, and determining an instruction type of the trigger instruction; A start unit consistent with the instruction type is matched from a type-start database, and the robot is started in a power-on state based on the start unit.
[0040] In this embodiment, the trigger instruction refers to whether the target user has a behavioral instruction to interact with the user, and the instruction type is a combination of any one or more of the voice interaction type, behavioral interaction type, and emotional interaction type.
[0041] In this embodiment, the type-startup database contains different instruction types and start-up units in the robot that match the type, that is, the start-up unit can trigger the start of the corresponding function of the robot, and then the robot is in the power-on state after startup.
[0042] The beneficial effect of the above technical solution is: by determining the instruction type, the starting unit is determined, and then the robot is started, and it is determined that subsequent data can be collected normally.
[0043] The present invention provides a multi-dimensional pixel fusion method based on multi-source perception, which controls various sensors set by the robot itself to collect data, including: When the robot is in a powered-on state, the robot collects an interactive input set of a target user currently received by the robot, and interactively analyzes the interactive input set to set a collection cycle for different types of sensors; The sensors of the corresponding type are controlled to collect data according to the collection period to obtain an initial set of each sensor.
[0044] In this embodiment, the interactive input set may be voice input information or text input information of the user.
[0045] In this embodiment, the purpose of interactive analysis is to determine the user's intention, and then set a collection cycle for the sensor to ensure the comprehensiveness and accuracy of the collection.
[0046] In this embodiment, the initial set includes collection results of the corresponding sensor in different collection cycles.
[0047] The beneficial effect of the above technical solution is: by analyzing the interactive input set to set the acquisition cycle, to determine the comprehensiveness and accuracy of the acquisition, and to provide a basis for subsequent pixel fusion.
[0048] The present invention provides a multi-dimensional pixel fusion method based on multi-source perception, which sets acquisition cycles for different types of sensors, including: Determining an input type of the interactive input set, and when the input type is voice input, respectively obtaining a response curve of each microphone collecting the user's voice; In combination with the setting position of the microphone of the robot itself and the current distance of the sound source of the target user, the response curves are sequentially sorted and aligned to obtain the response difference at each time point; Based on the response difference and time delay, a compensation adjustment is performed on the second curve which is the second closest, and waveform superposition processing is performed on the first curve which is the closest; Performing speech noise reduction processing on the superimposed speech to obtain the converted text; Extracting keywords from the converted text; Determine the duration of the process of the robot receiving the user's voice and performing semantic analysis on the user's voice, and at the same time, determine the voice duration of the user's voice; Obtaining a tone complexity coefficient according to the speech duration, the process duration, the tone information of the speech after the noise reduction process, and the number of words in the converted text; Retrieving a tone feature model that matches the tone complexity coefficient from a coefficient-feature comparison table to analyze the speech after noise reduction processing, and determining a tone annotation for each keyword; Merging the tone annotation with the extracted keywords to obtain a first annotation; When the input type is text input, extract keywords from the input text to obtain a second annotation; The interaction intention of the target user is determined according to the annotation, and the collection cycles of different types of sensors are matched from the intention comparison table.
[0049] In this embodiment, keyword extraction of text is achieved based on TF-IDF keyword extraction.
[0050] In this embodiment, the input type can only be a voice type or a text type.
[0051] In this embodiment, the robot itself is equipped with multiple microphones to ensure the reliability of the collected voice. However, due to the different locations of the microphones, in order to ensure the reliability of the voice, it is necessary to analyze the noise and other conditions in the voice collection process to make the voice more realistic.
[0052] In this embodiment, the response curve refers to a waveform curve of the user's voice collected by the corresponding microphone.
[0053] In this embodiment, since the positions of the microphones are slightly different, each acquired response curve is sorted in order according to the distance between the sound source and the microphone to determine the response difference at each time point: ,in, represents the energy difference between the first microphone and the second microphone at time point t after sequential sorting; represents the energy difference between the n-1th microphone and the nth microphone at time point t based on the sequential sorting; They respectively represent the audio energy of the 1st microphone, the 2nd microphone, the n-1th microphone, and the nth microphone at the sequentially sorted time point t.
[0054] In this embodiment, the time delay is obtained by matching from the position-delay comparison table, and different position distances correspond to different delay times, so that the time delay caused by the distance delay between the first microphone closest to the microphone and the second microphone closest to the microphone can be obtained, which facilitates the alignment of the two curves.
[0055] In this embodiment, the second curve which is closest to the first curve is aligned with the first curve after time shift according to the time delay, so as to obtain a third curve matching the second curve.
[0056] At this point, determine the delay factor for each time point based on the response difference: Each difference in the response difference at the corresponding time point is within the standard energy attenuation range at the corresponding distance between the two microphone positions. In this case, the attenuation factor is determined to be 0; Otherwise, it is determined that the attenuation factor is not 0, and the attenuation factor is calculated according to the following formula, that is: ; Where SJ represents the attenuation factor at the corresponding time point t; Nc represents the number of differences in the response difference at the corresponding time point t that are not within the standard energy attenuation range; M represents the total number of differences, and M=n-1; Indicates the absolute value of the difference that is not within the standard energy attenuation range for the i1th item; represents the maximum value in the i1th standard energy attenuation range; At this time, the energy value at the corresponding time point t in the third curve is adjusted according to SJ, that is: the energy A×ln(2+SJ) at the time point t in the third curve, thereby achieving compensation adjustment of the third curve.
[0057] In this embodiment, the speech noise reduction process may be implemented by using a filter.
[0058] In this embodiment, the semantic parsing model is obtained by training a neural network model based on different speech and semantic analysis results of the speech as samples. At this time, the duration of the semantic parsing process can be obtained by inputting the speech into the model, that is, the duration of the model analyzing the speech.
[0059] In this embodiment, the tone complexity coefficient = (process duration / speech duration) × (storage amount of tone information / storage amount of text words).
[0060] In this embodiment, the coefficient-feature comparison table includes different tone complexity coefficients and the corresponding tone feature models, and the larger the coefficient, the higher the accuracy of the corresponding tone feature model, which is mainly used to determine the tone annotation of keywords, such as sigh, surprise, astonishment, etc.
[0061] In this embodiment, the tone feature model is obtained by training a neural network model based on different keywords and the emotional tone results expressed by the keywords as samples. Therefore, the tone analysis and tone annotation can be directly implemented through the model, for example, keyword 1-surprise, keyword 2-surprise+anger.
[0062] In this embodiment, the first annotation includes all keywords and tone annotations for each keyword.
[0063] In this embodiment, when there are modal particles in the text, the corresponding keywords are directly annotated, and if not, no annotation is performed.
[0064] In this embodiment, the tone of voice can be used to assist in analyzing the user's true intention, thereby obtaining the true interaction intention, such as interacting with the robot through facial emotions, such as allowing the robot to imitate the user's facial expressions.
[0065] In this embodiment, the intent comparison table includes different interaction intentions and collection cycles that match the intentions. For example, if it is an emotion imitation intention, the collection cycle at this time is 0.1s. If it is a voice interaction intention, the collection cycle at this time is 1s. It should be noted that the collection cycles of different types of sensors may be the same or different.
[0066] The beneficial effect of the above technical solution is: the curves of the voice collected by the microphone are sorted in order to determine the corresponding differences at each time point, and then the curve closest to the second one is compensated and adjusted to ensure the reliability of the superposition processing with the second curve waveform, and then the text is obtained for keyword extraction and tone annotation, to ensure the authenticity of the interaction intention, provide an accurate basis for subsequent data collection, and ensure interaction efficiency.
[0067] The present invention provides a multi-dimensional pixel fusion method based on multi-source perception, aligns the extracted features under each type of sensor, and combines the acquisition array of the acquisition point and the tone attribute of the target user to perform multi-dimensional pixel fusion, including: Constructing a gradient vector for each pixel point in the acquisition array at each acquisition point, wherein the gradient vector includes a brightness gradient component of the corresponding pixel point under different types of sensors; Determine the information saturation coefficient of each pixel according to the gradient vector; When the information saturation coefficient is greater than a preset coefficient, the corresponding collection point is retained, otherwise, the corresponding collection point is removed; The extracted features of each type of sensor are converted into the spatial domain to obtain a spatial array of each spatial point. When there are remaining points that are consistent with the corresponding spatial points, the spatial array and the acquisition array under the consistent points are pixel-fused according to the tone attribute of the target user to obtain a fused array and retain it; When there are no corresponding spatial points that are consistent with the remaining points, the acquisition array corresponding to the remaining points is retained; Analyze the reserved array of each point to obtain behavioral operation instructions.
[0068] In this embodiment, the brightness gradient component is the image gradient, and the image gradient refers to the rate of change of a certain pixel of the image in the x and y directions (compared with adjacent pixels), that is, the gradient vector is a two-dimensional vector.
[0069] In this embodiment, the information saturation coefficient ,in, , They represent the rate of change based on the x and y directions respectively; r1, g1, l1 represent the actual pixel component values based on the red channel, green channel and blue channel respectively; r0, g0, l0 represent the maximum pixel component values based on the red channel, green channel and blue channel respectively; min represents the minimum value sign; ln represents the logarithmic function sign.
[0070] In this embodiment, the preset coefficient is set to 0.4.
[0071] In this embodiment, spatial domain conversion refers to performing feature-position mapping conversion on the extracted features, that is, comparing the features with the corresponding related position points one by one to match the features with the position points. For example, feature A1 corresponds to position 1 and position 2, that is, feature A1 is determined based on position 1 and position 2. It should be noted that the spatial points are related position points, and the number of spatial points is less than the number of position points.
[0072] In this embodiment, for example, the spatial points are u1, u2, u3, u4, u5, and the remaining points are u1, u2, u3, u4, u5, c1, c2, c3, c4, c5. At this time, the consistent points are: u1, u2, u3, u4, u5.
[0073] In this embodiment, the spatial array is the feature information of the corresponding spatial points under different types of sensors, and then the pixel information corresponding to the feature information is obtained from the feature-pixel comparison table, which contains different features and pre-set pixel conditions that match the features.
[0074] In this embodiment, the tone attribute is used to further optimize the array result to facilitate obtaining a fused array.
[0075] The beneficial effect of the above technical solution is: the information saturation coefficient of each pixel is calculated based on the gradient vector to preliminarily judge whether the collection point is retained, and then the spatial array is obtained by converting the features into the spatial domain and combining the collection array and the tone attribute to obtain a fusion array, so as to ensure the accuracy of the behavior operation instruction acquisition.
[0076] The present invention provides a multi-dimensional pixel fusion method based on multi-source perception, which performs pixel fusion on a spatial array and a collection array under a consistent point according to the tone attribute of the target user to obtain a fused array, including: Based on the spatial array and the acquisition array, construct an initial matrix; When the tone attribute is that the input type of the target user's interaction with the robot is unrelated to speech, an average value of each column element in the initial matrix is calculated to obtain a first array; When the tone attribute is an input type related to speech for the target user to interact with the robot, determining the facial area of the target user on which the consistent point acts according to the comparison relationship between the consistent point and the face of the target user, and setting simulation parameters for the consistent point according to the expression control database of the facial area; Extract the current tone of the target user at each time point during the interaction with the robot, and construct facial change information for each consistent point in combination with the user tone portrait of the target user and simulation parameters, and then construct a facial array based on the tone; The face array is appended to the first row of the initial matrix, and combined with the weight assigned to each row vector in the initial matrix, to obtain a second array; Among them, the first array and the second array are corresponding fused arrays.
[0077] Since the acquisition array = {pixel information collected by sensors of each type at the corresponding acquisition point}, the spatial array is the feature information of the corresponding spatial point under different types of sensors, and then the pixel information corresponding to the feature information is obtained from the feature-pixel comparison table.
[0078] Therefore, the initial matrix is = .
[0079] In this embodiment, the first array = [the average value of all pixel values in each column].
[0080] In this embodiment, the comparison relationship between the consistent point and the target user's face refers to the specific area where the point is located on the user's face, for example, the consistent point a1 is located in the apple muscle area of the left cheek of the user's face.
[0081] In this embodiment, the expression control database stores expression and emotion data related to different facial area blocks and simulation parameters for different expression and emotion data, and the simulation parameters are set to determine the parameters of the robot's facial emotion simulation, and the simulation parameters corresponding to different facial expressions are different. For example, the simulation parameters corresponding to the facial expression of laughing are: a1, a2, a4, a6, and laughter is achieved through simulation.
[0082] In this embodiment, the user tone portrait is pre-set, and the user tone portrait refers to the facial expression expressed by the user's tone, and the user tone portrait can be obtained by performing a portrait analysis on the user's facial expressions under different tones, which belongs to the prior art.
[0083] In this embodiment, the current tone, user tone portrait, and simulation parameters are input into the facial analysis model to obtain corresponding facial change information. The facial analysis model is based on different tones, user tone portraits, and simulation parameters, and the facial change results for these parameters are obtained by training the neural network model with samples. Therefore, facial change information can be directly obtained, such as facial changes of the consistent point under the corresponding tone, such as the cheek position C1 changes from height h1 to height h2.
[0084] Face array = {pixel information of different sensor types based on facial change information at each consistent point}. It should be noted that the pixel information of facial change information under different sensor types is obtained by matching from the transformation-pixel comparison table. The comparison table contains the pixel conversion of the final change result of the facial change information under different sensor types, and thus the corresponding pixel information can be directly obtained.
[0085] In this embodiment, the weights are pre-set. Generally, the weights corresponding to the facial array, acquisition array, and spatial array are 0.3, 0.4, and 0.3, respectively. Then, the pixel information of the corresponding column is: weight × the sum of the element values of the corresponding column, and the second array = {the sum of the element values of each column}.
[0086] The beneficial effect of the above technical solution is: constructing an initial matrix based on the spatial array and the acquisition array, and then analyzing the arrays in different situations based on whether the voice is relevant. When it is related to the voice, based on the tone and portrait, and combining the expression control database to set the simulation parameters, construct a facial array to optimize the initial matrix and achieve accurate acquisition of the fusion array.
[0087] The present invention provides a multi-dimensional pixel fusion method based on multi-source perception, which analyzes the reserved array of each point to obtain a behavior operation instruction, including: Divide the retained array of each point based on the block division criteria of the user's face to obtain the current representation of each block; Based on all current representations, the block weight of each divided block and the user intention determined by the user voice of the target user, a behavior operation instruction is obtained from the intention-combination-instruction comparison table.
[0088] In this embodiment, the user's face is divided into blocks based on the following criteria: nose, eyebrows, eyes, forehead, chin, cheeks, etc.
[0089] In this embodiment, the points belonging to the divided blocks are assigned to the corresponding blocks to obtain all the retained arrays of each divided block, and all the retained arrays under each divided block are sequentially input into the corresponding facial area model to obtain the current representation. The facial area model is obtained by training the neural network model based on the pixel arrays under different expressions of the area and the behavioral representation of the user's real emotions as samples. Therefore, the current representation can be directly obtained.
[0090] In this embodiment, the weight of each divided block is preset, and the sum of the weights is 1. It should be noted that the weights corresponding to different facial regions may be different.
[0091] In this embodiment, the intention-combination-instruction comparison table includes the intention, the current representation of different combinations, the block weights and the matching behavioral operation instructions, and the behavioral operation instructions can be directly obtained. It should be noted that the operation instructions can be directly determined based on the intention, but through further analysis of the face and tone, the accuracy of the instruction determination can be further guaranteed.
[0092] The beneficial effect of the above technical solution is: by determining the current representation of each facial area and combining the user intention and block weight, the behavioral operation instructions are directly obtained, ensuring the accuracy of instruction acquisition, thereby improving the interaction efficiency.
[0093] The present invention provides a multi-dimensional pixel fusion system based on multi-source perception, such as Figure 2 As shown, including: A data acquisition module, used to control various sensors set by the robot itself to collect data, wherein the sensors include: a visible light camera, an infrared camera and a laser radar; The alignment processing module is used to unify the coordinates and align the timing of the collected data of each sensor to obtain the collection array of each collection point; A feature extraction module is used to arrange and process the collected data of each type of sensor in spatial order to obtain a corresponding collected image, and perform feature extraction on the collected image according to a feature extraction method matching the corresponding type of sensor; The multi-dimensional pixel fusion module is used to align the extracted features under each type of sensor, and perform multi-dimensional pixel fusion in combination with the acquisition array of the acquisition points and the tone attributes of the target user, obtain behavioral operation instructions and output them to the control unit of the robot to perform corresponding operations.
[0094] The beneficial effect of the above technical solution is: starting from multi-source perception to determine multi-dimensional pixel information, and combining the influence of tone feedback in the user's voice process to ensure the reliability of multi-dimensional pixel fusion, thereby making behavioral operation instructions more accurate, to improve interaction efficiency and satisfy the user's interactive experience.
[0095] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A multi-dimensional pixel fusion method based on multi-source perception, characterized in that: include: Step 1: Control the sensors set on the robot to collect data, wherein the sensors include: a visible light camera, an infrared camera and a laser radar; Step 2: Unify the coordinates and align the time sequence of the collected data of each sensor to obtain the collection array of each collection point; Step 3: Arrange and process the collected data of each type of sensor in spatial order to obtain a corresponding collected image, and perform feature extraction on the collected image according to a feature extraction method matching the corresponding type of sensor; Step 4: Align the extracted features under each type of sensor, and perform multi-dimensional pixel fusion in combination with the acquisition array of the acquisition point and the tone attribute of the target user to obtain behavioral operation instructions and output them to the control unit of the robot to perform corresponding operations.
2. The multi-dimensional pixel fusion method based on multi-source perception according to claim 1, characterized in that: Before controlling the sensors set on the robot to collect data, it includes: Capturing a trigger instruction of the robot from external information, and determining an instruction type of the trigger instruction; A start unit consistent with the instruction type is matched from a type-start database, and the robot is started in a power-on state based on the start unit.
3. The multi-dimensional pixel fusion method based on multi-source perception according to claim 2, characterized in that: Control the sensors set on the robot to collect data, including: When the robot is in a powered-on state, the robot collects an interactive input set of a target user currently received by the robot, and interactively analyzes the interactive input set to set a collection cycle for different types of sensors; The sensors of the corresponding type are controlled to collect data according to the collection period to obtain an initial set of each sensor.
4. The multi-dimensional pixel fusion method based on multi-source perception according to claim 3 is characterized in that: Set the collection period for different types of sensors, including: Determining an input type of the interactive input set, and when the input type is voice input, respectively obtaining a response curve of each microphone collecting the user's voice; In combination with the setting position of the microphone of the robot itself and the current distance of the sound source of the target user, the response curves are sequentially sorted and aligned to obtain the response difference at each time point; Based on the response difference and time delay, a compensation adjustment is performed on the second curve which is the second closest, and waveform superposition processing is performed on the first curve which is the closest; Performing speech noise reduction processing on the superimposed speech to obtain the converted text; Extracting keywords from the converted text; Determine the duration of the process of the robot receiving the user's voice and performing semantic analysis on the user's voice, and at the same time, determine the voice duration of the user's voice; Obtaining a tone complexity coefficient according to the speech duration, the process duration, the tone information of the speech after the noise reduction process, and the number of words in the converted text; Retrieving a tone feature model that matches the tone complexity coefficient from a coefficient-feature comparison table to analyze the speech after noise reduction processing, and determining a tone annotation for each keyword; Merging the tone annotation with the extracted keywords to obtain a first annotation; When the input type is text input, extract keywords from the input text to obtain a second annotation; The interaction intention of the target user is determined according to the annotation, and the collection cycles of different types of sensors are matched from the intention comparison table.
5. The multi-dimensional pixel fusion method based on multi-source perception according to claim 1, characterized in that: The extracted features of each type of sensor are aligned, and multi-dimensional pixel fusion is performed in combination with the acquisition array of the acquisition point and the tone attribute of the target user, including: Constructing a gradient vector for each pixel point in the acquisition array at each acquisition point, wherein the gradient vector includes a brightness gradient component of the corresponding pixel point under different types of sensors; Determine the information saturation coefficient of each pixel according to the gradient vector; When the information saturation coefficient is greater than a preset coefficient, the corresponding collection point is retained, otherwise, the corresponding collection point is removed; The extracted features of each type of sensor are converted into the spatial domain to obtain a spatial array of each spatial point. When there are remaining points that are consistent with the corresponding spatial points, the spatial array and the acquisition array under the consistent points are pixel-fused according to the tone attribute of the target user to obtain a fused array and retain it; When there are no corresponding spatial points that are consistent with the remaining points, the acquisition array corresponding to the remaining points is retained; Analyze the reserved array of each point to obtain behavioral operation instructions.
6. The multi-dimensional pixel fusion method based on multi-source perception according to claim 5, characterized in that: The spatial array and the collection array at the consistent point are pixel-fused according to the tone attribute of the target user to obtain a fused array, including: Based on the spatial array and the acquisition array, construct an initial matrix; When the tone attribute is that the input type of the target user's interaction with the robot is unrelated to speech, an average value of each column element in the initial matrix is calculated to obtain a first array; When the tone attribute is an input type related to speech for the target user to interact with the robot, determining the facial area of the target user on which the consistent point acts according to the comparison relationship between the consistent point and the face of the target user, and setting simulation parameters for the consistent point according to the expression control database of the facial area; Extract the current tone of the target user at each time point during the interaction with the robot, and construct facial change information for each consistent point in combination with the user tone portrait of the target user and simulation parameters, and then construct a facial array based on the tone; The face array is appended to the first row of the initial matrix, and combined with the weight assigned to each row vector in the initial matrix, to obtain a second array; Among them, the first array and the second array are corresponding fused arrays.
7. The multi-dimensional pixel fusion method based on multi-source perception according to claim 6, characterized in that: Analyze the reserved array of each point to obtain behavioral operation instructions, including: Divide the retained array of each point based on the block division criteria of the user's face to obtain the current representation of each block; Based on all current representations, the block weight of each divided block and the user intention determined by the user voice of the target user, a behavior operation instruction is obtained from the intention-combination-instruction comparison table.
8. A multi-dimensional pixel fusion system based on multi-source perception, characterized in that: include: A data acquisition module, used to control various sensors set by the robot itself to collect data, wherein the sensors include: a visible light camera, an infrared camera and a laser radar; The alignment processing module is used to unify the coordinates and align the timing of the collected data of each sensor to obtain the collection array of each collection point; A feature extraction module is used to arrange and process the collected data of each type of sensor in spatial order to obtain a corresponding collected image, and perform feature extraction on the collected image according to a feature extraction method matching the corresponding type of sensor; The multi-dimensional pixel fusion module is used to align the extracted features under each type of sensor, and perform multi-dimensional pixel fusion in combination with the acquisition array of the acquisition points and the tone attributes of the target user, obtain behavioral operation instructions and output them to the control unit of the robot to perform corresponding operations.
Citation Information
Patent Citations
Intelligent interaction method and device, computer equipment and computer readable storage medium
CN108197115A
Control system and method of artificial intelligence drive robot
CN108297098A
Robot-oriented multimodal fusion emotion calculation method and system
CN108960191A
Intelligent customer service system based on AI large model
CN119474280A
Multimodal based interaction robot, and control method for the same
KR102519599B1