A desktop companion robot audio-visual composite positioning and orientation control method and system
Patent Information
- Application Number
- CN202611081529.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-21
- Publication Date
- 2026-08-18
AI Technical Summary
针对现有技术的不足,本发明提供了一种桌面陪护机器人声视复合定位朝向控制方法及系统,解决了现有机器人在声源定位与视觉复定位后,难以根据头部、颈部和躯体的转动范围及响应差异分配朝向动作,导致摄像头采集方向、显示屏展示方向和语音交互方向难以同步对准用户的问题
(1)本发明,通过声源定位、视觉复定位和机器人多自由度朝向控制的协同处理,使桌面陪护机器人能够在确认用户方位后对头部、颈部和躯体进行分工控制,进而实现了机器人多自由度机构协同面向用户的效果,有效解决了现有技术中桌面陪护机器人仅依据用户方位生成单一转向动作、难以协调多个执行机构的问题。
Smart Images

Figure CN122593432A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of posture control technology, specifically to a method and system for acoustic-visual composite positioning and orientation control of a desktop companion robot. Background Technology
[0002] With the aging population and the increasing demand for home-based elderly care, intelligent companionship devices for family companionship, health reminders, daily information services, and remote communication are gradually being applied. Desktop companion robots, due to their small size, easy deployment, and short interaction distance, can be placed in living rooms, bedrooms, bedside tables, desktops, and other scenarios to provide users with services such as voice communication, information broadcasting, facial expression display, motion feedback, and remote interaction with relatives. Existing desktop companion robots typically integrate display, voice, image acquisition, network communication, and motion execution functions. They display facial expressions or information through a screen, receive user voice through a microphone, assist user perception through a camera, and enhance interactive performance through head, neck, or body movements. As companion robots evolve from simple voice response devices to intelligent terminals with facial expressions, motion, and scene linkage capabilities, robots need to not only complete voice broadcasting and information display when communicating with users, but also need to coordinate with corresponding postures and movements to improve the naturalness, friendliness, and sense of companionship in the interaction.
[0003] For example, the invention patent with announcement number CN107199572B discloses a robot system and method based on intelligent sound source localization and voice control. The robot continuously collects surrounding voice information. When a voice command is received, it performs sound source localization, controls the robot to move to the sound source location, and recognizes the collected voice information. When a valid sentence is recognized, the corresponding control command is sent to the robot to perform the corresponding operation. At the same time, the valid sentence is translated into corresponding text, Chinese word segmentation is performed, and a sentiment dictionary, degree adverb dictionary, negation word list, and association word list are loaded to identify each sentiment word in the sentence. Based on the recognition results, the corresponding expression is displayed, realizing robot interactive control based on sound source localization, voice recognition, and emotion recognition, and improving the interaction ability between service robots and those being accompanied.
[0004] For example, the invention patent with announcement number CN111055288B discloses a robot control method, storage medium, and robot that can respond to calls on demand. After obtaining the sound source localization direction information, the robot obtains a list of search target points and navigates to each search target point in sequence according to the list. During navigation, it obtains information about surrounding objects and analyzes whether there are people in each one. When a person is detected, it determines whether the detected person is the target person. If it is the target person, it calculates the distance from the robot to the target person and navigates to a stop point at a preset distance from the target person. If it is not the target person, it returns to continue navigation and detection. This realizes the function of making the robot respond to calls on demand through voice recognition, sound source localization direction, human body detection, and face recognition technologies, and improves the interactivity and human-like characteristics of the robot.
[0005] However, existing desktop companion robots mostly focus on determining the user's location or the identity of the target person through sound source localization, speech recognition, human detection, or facial recognition, and then performing interactive actions such as movement, turning, voice response, or facial expression display. For the orientation control process after the user's location has been determined, there is usually a lack of coordinated allocation of the rotation range, response speed, and limit states among multiple actuators of the head, neck, and torso. Especially when the desktop companion robot needs to simultaneously consider the camera's acquisition direction, the display screen's display direction, and the voice interaction direction, generating a single turning action based solely on the user's location can easily lead to local mechanisms exceeding their reach, redundant head and torso compensation, or the display screen failing to effectively face the user, thus affecting the robot's stability and the continuity of interaction.
[0006] Therefore, in order to address the above problems, there is an urgent need for a method and system for controlling the orientation of a desktop companion robot using a combination of sound and vision. Summary of the Invention
[0007] Technical problems to be solved To address the shortcomings of existing technologies, this invention provides a method and system for controlling the orientation of a desktop companion robot using a combination of sound and vision positioning. This solves the problem that existing robots, after sound source localization and visual relocalization, struggle to allocate orientation movements based on the rotation range and response differences of the head, neck, and torso, resulting in the camera acquisition direction, display screen direction, and voice interaction direction being difficult to align synchronously with the user.
[0008] Technical solution To achieve the above objectives, the present invention provides the following technical solution: a method for acoustic-visual composite positioning and orientation control of a desktop companion robot, comprising: S1, collecting raw data across sensors and performing time alignment, data correction, anomaly removal, and standardization to form an acoustic-visual attitude synchronization data chain; S2, generating candidate sectors of sound sources and extracting visual target features based on the acoustic-visual attitude synchronization data chain, and obtaining a user orientation confidence value through evidence fusion; S3, constructing a head, neck, and body orientation angle constraint set to obtain the orientation deviation to be eliminated, and generating a temporary acceptance sequence, orientation execution sequence, line-of-sight correction sequence, or limit avoidance mark based on the head, neck, and body orientation angle constraint set and the orientation deviation to be eliminated; S4, generating motor control commands based on the temporary acceptance sequence, orientation execution sequence, or line-of-sight correction sequence, executing orientation control, and outputting an orientation completion mark.
[0009] Furthermore, the specific process for collecting cross-sensor raw data is as follows: A desktop companion robot equipped with a head display screen, head camera, left microphone, right microphone, neck pitch mechanism, body rotation mechanism, and motor drive board is used as the data acquisition object. Cross-sensor raw data is collected, including: synchronously acquiring the left microphone voice sampling sequence, right microphone voice sampling sequence, and acoustic sampling timestamp through the left and right microphones; acquiring camera image frame sequences, image sampling timestamps, image width, and image height through the head camera; acquiring angle sequences through the head, neck, and body angle encoders, including: head yaw angle sequence, neck pitch angle sequence, and body rotation angle sequence; during the robot's self-test actions, acquiring limit angle data through limit switches and angle encoders; and acquiring motor control command timestamps and motor feedback current sequences through the motor drive board, including: head motor feedback current sequence, neck motor feedback current sequence, and body motor feedback current sequence.
[0010] Further, the specific process of forming an audio-visual attitude synchronization data chain by performing time alignment, data correction, anomaly removal, and normalization is as follows: The left microphone speech sampling sequence, right microphone speech sampling sequence, and camera image frame sequence are arranged on the same interactive time axis according to the acoustic sampling timestamp and image sampling timestamp; silent segment clipping, clipped segment removal, and left and right channel amplitude correction are performed on the left and right microphone speech sampling sequences; jitter frame removal, exposure consistency, and distortion correction are performed on the camera image frame sequence; jump points are removed from the angle sequence and motor feedback current sequence to form an audio-visual attitude synchronization data chain; the left microphone speech sampling sequence, right microphone speech sampling sequence, and motor feedback current sequence are normalized using the Z-score normalization algorithm; and the pixel brightness values and angle sequences in the camera image frame sequence are normalized using the Min-Max normalization algorithm.
[0011] Furthermore, the specific process of generating candidate sectors for sound sources and extracting visual target features based on the audio-visual attitude synchronization data link is as follows: extracting wake-up speech segments from the audio-visual attitude synchronization data link; dividing the left microphone speech sampling sequence and the right microphone speech sampling sequence into multiple speech frames according to a fixed frame length; extracting the amplitude envelopes of the left and right channels within each speech frame; taking the time corresponding to the maximum point of the amplitude envelope as the arrival time of the left channel peak and the arrival time of the right channel peak, and subtracting the arrival time of the left channel peak from the arrival time of the right channel peak to obtain the two-channel arrival time difference sequence; determining the left and right directions of the sound source based on the positive and negative distribution of the two-channel arrival time difference sequence; determining the lateral angle range corresponding to the candidate sectors of the sound source based on the maximum and minimum values of the two-channel arrival time difference sequence; when the two-channel arrival time difference sequence in consecutive speech frames... When the positive and negative signs of the column change alternately, the azimuth range corresponding to the alternating speech frames is incorporated into the sound source candidate sector, and direct body turning is paused. Body turning is resumed after the current visual candidate target is confirmed as the current interactive target. Face candidate boxes are extracted within the camera image frame region corresponding to the sound source candidate sector. The center coordinates are obtained from the coordinates of the upper left and lower right corners of the face candidate box to obtain the center point of the face candidate box. The horizontal and vertical coordinates of the center of the face candidate box are extracted. The lower half of the face candidate box is taken as the mouth region. The Kanade-Lucas-Tomasi optical flow tracking algorithm is used to track the feature points of the mouth region to obtain the mouth motion trajectory. The trajectory of the visual candidate target is formed according to the displacement change of the center point of the face candidate box in continuous image frames.
[0012] Furthermore, the specific process of obtaining the user's location confidence value through evidence fusion is as follows: Half the image width is subtracted from the horizontal coordinate of the face candidate box center to obtain the user's horizontal deviation pixels; half the image height is subtracted from the vertical coordinate of the face candidate box center to obtain the user's vertical deviation pixels; the proportion of frames in the camera image frame region corresponding to the sound source candidate sector where the face candidate box center is located is calculated to obtain the sound source continuity score; the proportion of the overlap duration between the image frame time window corresponding to the peak of the mouth movement trajectory and the wake-up speech segment is calculated to obtain the mouth movement synchronization score; the proportion of the number of consecutively existing frames of the visual candidate target trajectory is calculated to obtain the trajectory continuity score; the sound source continuity score, mouth movement synchronization score, and trajectory continuity score are fused using the Dempster-Shafer-based evidence discount fusion method to obtain the user's location confidence value; when the user's location confidence value is less than the user's location confidence threshold, a visual verification mark is generated; when the user's location confidence value is greater than or equal to the user's location confidence threshold, the current visual candidate target is confirmed as the current interaction target.
[0013] Furthermore, the specific process of constructing a set of head, neck, and body orientation angle constraints to obtain the orientation deviation to be eliminated is as follows: Read the user's lateral deviation pixels, user's longitudinal deviation pixels, user's orientation confidence value, and visual verification markers; determine the left and right deviation directions based on the sign of the user's lateral deviation pixels, and determine the pitch deviation direction based on the sign of the user's longitudinal deviation pixels; use the current sampling angle in the angle sequence as the attitude reference; use the corresponding limit angle data as the motion boundary, and trim the head yaw angle adjustment range, neck pitch angle adjustment range, and body rotation angle adjustment range according to the deviation direction; extract the image center region from the camera image frame sequence; determine the camera's field of view coverage area based on the range where the center of the face candidate box falls within the image center region. The display screen's forward orientation angle range is determined by the inclusion relationship between the torso rotation angle and the user's lateral angle range, forming a set of head, neck, and body orientation angle constraints. Based on the motor control command timestamp, single-degree-of-freedom motion segments with only a single motor action and other angle changes within the angle sampling fluctuation range are selected. The displacement of the face candidate box center before and after the single-degree-of-freedom motion segment, as well as the difference in the corresponding head yaw angle, neck pitch angle, or torso rotation angle before and after the action, are calculated to form the pixel angle response relationship of the head, neck, and torso. Then, based on the pixel angle response relationship, the pixel angles of the user's lateral deviation pixels and user's longitudinal deviation pixels are inversely calculated to obtain the orientation deviation to be eliminated.
[0014] Furthermore, the specific process of generating temporary acceptance sequences, orientation execution sequences, line-of-sight correction sequences, or limit avoidance markers based on the head, neck, and body orientation angle constraint set and the orientation deviation to be eliminated is as follows: Within the action execution window, the time difference from the motor control command timestamp to entering the drive current segment and the time difference from the drive current segment back to the stationary current segment are calculated based on the motor feedback current sequence, and the two types of time differences are added to obtain the head response duration, neck response duration, and body response duration; Based on the head, neck, and body orientation angle constraint set and the orientation deviation to be eliminated, a hierarchical model predictive control process with limit constraints is established: The first layer aims to recapture the face candidate box by the camera, compares the head response duration and neck response duration, and determines whether the head yaw angle adjustment range and neck pitch angle adjustment range cover the corresponding orientation deviation to be eliminated. Among the heads or necks that cover the corresponding orientation deviation to be eliminated at the turning angle, the mechanism corresponding to the smaller value of the head response duration and neck response duration is selected to perform the line-of-sight acceptance; The second layer aims to face the user with the display screen, and calculates the time difference from the motor control command timestamp to entering the drive current segment and the time difference from the drive current segment back to the stationary current segment, and adds the two types of time differences to obtain the head response duration, neck response duration, and body response duration. Within the adjustment range, select the body rotation direction with the smaller absolute value of the rotation angle between clockwise and counterclockwise directions; the third layer aims to reduce redundant compensation, ensuring that the head correction angle only compensates for the remaining lateral deviation components not covered by the main body steering angle; when a visual verification mark exists, generate the head yaw acceptance angle and neck pitch acceptance angle, and output a temporary acceptance sequence; when no visual verification mark exists and the display screen's forward display area does not cover the user's lateral angle range calculated from the user's lateral deviation pixels, generate the main body steering angle, head correction angle, and neck compensation angle, and output an orientation execution sequence; when no visual verification mark exists and the display screen's forward display area covers the user's lateral angle range calculated from the user's lateral deviation pixels, generate the head correction angle and neck compensation angle, and output a line-of-sight correction sequence; when the head yaw acceptance angle or head correction angle exceeds the head yaw angle adjustment range, the neck pitch acceptance angle or neck compensation angle exceeds the neck pitch angle adjustment range, or the main body steering angle exceeds the body rotation angle adjustment range, output a limit avoidance mark.
[0015] Furthermore, the specific process of generating motor control commands based on the temporary acceptance sequence, orientation execution sequence, or line-of-sight correction sequence is as follows: According to the temporary acceptance sequence, orientation execution sequence, or line-of-sight correction sequence, the yaw angle allocated to the head, the pitch angle allocated to the neck, and the rotation angle allocated to the torso are converted into head motor control commands, neck motor control commands, and torso motor control commands, respectively, and sent to the motor drive board; when there is a limit avoidance mark, a substitute acceptance angle is generated based on the remaining rotation angle of the mechanism that does not exceed the adjustment range of the head yaw angle, the adjustment range of the neck pitch angle, or the adjustment range of the torso rotation angle, and the execution angle that exceeds the range is replaced by the substitute acceptance angle, so that the face candidate box enters the camera's field of view acceptance range; when there is no limit avoidance mark, the corresponding mechanism action is executed in the order of the temporary acceptance sequence, orientation execution sequence, or line-of-sight correction sequence.
[0016] Furthermore, the specific process of executing orientation control and outputting the orientation completion mark is as follows: during the execution process, the angle sequence, the motor feedback current sequence, and the camera image frame sequence are continuously read; the target angle in the motor control command is compared with the final sampled angle of the angle sequence after the corresponding control command is executed, and the absolute difference is calculated to obtain the head execution deviation, neck execution deviation, and body execution deviation, and the user orientation confidence value is recalculated. Orientation control process: When the updated user orientation confidence value is greater than or equal to the user orientation confidence threshold, maintain the current head yaw angle, neck pitch angle, and body rotation angle; when the updated user orientation confidence value is less than the user orientation confidence threshold, sequentially execute head left boundary pause frame acquisition, head middle position pause frame acquisition, and head right boundary pause frame acquisition within the sound source candidate sector; if the recalculated user orientation confidence value is still less than the user orientation confidence threshold, execute neck downward pause frame acquisition and neck upward pause frame acquisition; if the lateral angle range corresponding to the sound source candidate sector exceeds the head yaw angle adjustment range, control the body to perform segmented compensation rotation along the direction of the sound source candidate sector, and reconfirm the current interaction target; after the orientation control process is completed, output orientation completion mark, orientation deviation record, degree of freedom limit record, and interaction acceptance status.
[0017] The second aspect of this invention provides a desktop companion robot's audio-visual composite positioning and orientation control system, comprising: an acquisition and preprocessing unit, used to acquire raw data across sensors and perform time alignment, data correction, anomaly removal, and standardization to form an audio-visual attitude synchronization data chain; an audio-visual composite orientation receiving unit, used to generate candidate sectors of sound sources and extract visual target features based on the audio-visual attitude synchronization data chain, and obtain a user orientation confidence value through evidence fusion; a multi-degree-of-freedom orientation allocation unit, used to construct a head, neck, and body orientation angle constraint set, obtain the orientation deviation to be eliminated, and generate a temporary receiving sequence, an orientation execution sequence, a line-of-sight correction sequence, or a limit avoidance mark based on the head, neck, and body orientation angle constraint set and the orientation deviation to be eliminated; and a segmented execution and closed-loop correction unit, used to generate motor control commands based on the temporary receiving sequence, the orientation execution sequence, or the line-of-sight correction sequence, execute orientation control, and output an orientation completion mark.
[0018] Beneficial effects The present invention has the following beneficial effects: (1) This invention enables the desktop companion robot to perform division of labor control on the head, neck and torso after confirming the user's position through the coordinated processing of sound source localization, visual relocalization and robot multi-degree-of-freedom orientation control, thereby realizing the effect of robot multi-degree-of-freedom mechanism coordinating to face the user, effectively solving the problem in the prior art that the desktop companion robot only generates a single turning action based on the user's position and it is difficult to coordinate multiple actuators.
[0019] (2) This invention achieves the effect of mutual verification between acoustic perception and visual perception by combining the candidate sector of the sound source with the features of the visual target and combining the user's location confidence value to confirm the current interactive target. This effectively solves the problem that relying solely on sound source localization or visual recognition in the prior art can easily lead to misjudgment of the interactive target.
[0020] (3) This invention constructs a set of head, neck and body orientation angle constraints, incorporates the rotation boundaries of the head, neck and body into the orientation allocation process, thereby achieving the effect of allocating orientation movements within the reach of the mechanism, effectively solving the problem in the prior art where local degrees of freedom easily reach the limit and cause the orientation adjustment to be interrupted.
[0021] (4) The present invention converts the user’s deviation in the image into a robot-executable orientation deviation through the pixel angle response relationship, and further generates a temporary receiving sequence, orientation execution sequence or gaze correction sequence, thereby realizing the orderly mapping effect from visual deviation to mechanism action, effectively solving the problem of inaccurate connection between image target position and robot execution action in the prior art.
[0022] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description
[0023] Figure 1 This is a flowchart of a desktop companion robot's acoustic-visual composite positioning and orientation control method according to the present invention; Figure 2 This is a structural diagram of a desktop companion robot's audio-visual composite positioning and orientation control system according to the present invention; Figure 3 This is a scatter plot of the user deviation pixels and orientation allocation results of the present invention; Figure 4 This is a distribution diagram of the user's location confidence value and lateral deviation in this invention; Figure 5 This is a flowchart of the multi-degree-of-freedom motor command conversion and limit avoidance execution of the present invention. Detailed Implementation
[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] Please see Figures 1-5This invention provides a technical solution: a method for acoustic-visual composite positioning and orientation control of a desktop companion robot, comprising the following steps: S1, collecting raw data across sensors and performing time alignment, data correction, anomaly removal, and standardization to form an acoustic-visual attitude synchronization data chain; S2, generating candidate sectors for sound sources and extracting visual target features based on the acoustic-visual attitude synchronization data chain, and obtaining a user orientation confidence value through evidence fusion; S3, constructing a head, neck, and body orientation angle constraint set to obtain the orientation deviation to be eliminated, and generating a temporary acceptance sequence, orientation execution sequence, line-of-sight correction sequence, or limit avoidance mark based on the head, neck, and body orientation angle constraint set and the orientation deviation to be eliminated; S4, generating motor control commands based on the temporary acceptance sequence, orientation execution sequence, or line-of-sight correction sequence, executing orientation control, and outputting an orientation completion mark.
[0026] Specifically, the process of collecting cross-sensor raw data is as follows: A desktop companion robot equipped with a head-mounted display, head-mounted camera, left microphone, right microphone, neck pitch mechanism, body rotation mechanism, and motor drive board is used as the data acquisition object. Its multimodal sensors are triggered by a unified clock source and synchronized with hardware timestamps. Cross-sensor raw data is collected, including: synchronously acquiring the left microphone voice sampling sequence, right microphone voice sampling sequence, and acoustic sampling timestamp through the left and right microphones. The left and right microphones are synchronously started for dual-channel sampling via an audio codec using frame interrupts; acquiring camera image frame sequences, image sampling timestamps, image width, and image height through the head-mounted camera. Subsequent data involving camera image frame sequences, image sampling timestamps, image width, image height, and the camera's field of view all originate from the head-mounted camera. The camera module outputs a frame synchronization signal with a hardware timestamp via the MIPI interface; through the head-mounted camera... Angle encoders for the neck and torso collect angle sequences, including head yaw angle, neck pitch angle, and torso rotation angle. The angle encoders use magnetic Hall elements for quadrature decoding to generate incremental angle counts. During robot self-testing, limit switches and angle encoders collect limit angle data, including left head limit angle, right head limit angle, neck tilt limit angle, neck downward limit angle, torso clockwise limit angle, and torso counterclockwise limit angle, corresponding to the two-way movement boundaries of head yaw, neck pitch, and torso rotation, respectively. When a limit switch is triggered, the real-time reading of the corresponding angle encoder is locked and recorded as limit angle data. The motor drive board collects motor control command timestamps and motor feedback current sequences, including head motor feedback current sequences, neck motor feedback current sequences, and torso motor feedback current sequences. The motor drive board continuously collects coil currents through sampling resistors and an analog-to-digital converter and adds command response timestamps.
[0027] In this implementation plan, this step achieves unified collection of acoustic, visual, posture, limit, and drive feedback data of the desktop companion robot, forming a raw data foundation covering user voice triggering, face image observation, head, neck, and body posture status, mechanism motion boundaries, and motor response status. This provides complete, synchronous, and traceable data support for subsequent construction of an audio-visual posture synchronization data chain, determination of user orientation reliability values, and allocation of head yaw acceptance angle, neck pitch acceptance angle, and body main steering angle.
[0028] Specifically, the process of forming a synchronized audio-visual attitude data chain through time alignment, data correction, anomaly removal, and standardization is as follows: Based on the acoustic sampling timestamps and image sampling timestamps, and using the hardware timestamp under a unified clock source as the time reference, the left microphone speech sampling sequence, right microphone speech sampling sequence, and camera image frame sequence are arranged on the same interactive time axis. Cross-modal sub-frame alignment is achieved through nearest neighbor timestamp matching and linear interpolation resampling. Nearest neighbor timestamp matching is used to determine the correspondence between acoustic sampling timestamps and image sampling timestamps, and linear interpolation resampling is used to fill in the sampling differences between the correspondences. The left and right microphone speech sampling sequences undergo silent segment clipping, clipped segment removal, and left and right channel amplitude correction. Silent segment clipping is based on a dual-threshold decision of short-time energy and zero-crossing rate. Clipped segment removal uses neighborhood spline interpolation repair of peak cutoff samples. Left and right channel amplitude correction is compensated by calculating the root mean square ratio between channels. The camera image frame sequence undergoes jitter frame removal and exposure correction. For image homogenization and distortion correction, jittery frame removal uses inter-frame structural similarity exponential mutation detection, exposure homogenization uses grayscale histogram specification mapping to reference frames, and distortion correction is based on the camera calibration intrinsic parameter matrix and radial-tangential distortion coefficient polynomial model. Jump point removal is performed on angle sequences and motor feedback current sequences to form a sound-visual attitude synchronization data chain. Jump point removal uses sliding window mid-range filtering and differential absolute value thresholding to identify isolated outliers. The left microphone speech sampling sequence, right microphone speech sampling sequence, and motor feedback current sequence are standardized using the Z-score normalization algorithm to convert each sequence into a distribution with a mean of zero and a standard deviation of one to eliminate differences in physical dimensions. The pixel brightness values and angle sequences in the camera image frame sequence are normalized using the Min-Max normalization algorithm to linearly map the original values to the closed interval [0, 1] to retain relative proportion features. Here, the Z-score normalization algorithm represents the Z-score normalization algorithm, and the Min-Max normalization algorithm represents the maximum-minimum normalization algorithm.
[0029] This implementation plan completes the unified temporal processing, effective signal preservation, abnormal interference removal, and feature scale consistency processing of acoustic data, visual data, posture data, and drive feedback data. This forms a synchronous audio-visual posture data chain with consistent time reference, stable data quality, and unified parameter expression. It reduces the impact of speech segment noise, image frame jitter, posture jumps, and current anomalies on subsequent analysis results. It provides a continuous, reliable, and comparable data foundation for the generation of candidate sectors for sound sources, extraction of visual target features, calculation of user orientation confidence values, construction of head, neck, and body orientation angle constraint sets, and back-calculation of orientation deviations to be eliminated.
[0030] Specifically, the process of generating candidate sectors for sound sources and extracting visual target features based on the audio-visual attitude synchronization data link is as follows: A wake-up speech segment is extracted from the audio-visual attitude synchronization data link. The wake-up speech segment is determined by detecting the start and end points in the continuous audio stream using a pre-set keyword wake-up engine, and the audio segment between the start and end points is taken as the wake-up speech segment. The left microphone speech sampling sequence and the right microphone speech sampling sequence are divided into multiple speech frames with a fixed frame length. The fixed frame length is calculated from the acoustic sampling frequency obtained from the acoustic sampling timestamp, and is between 20ms and 40ms, ensuring that the short-time energy changes within a single speech frame meet the requirements for short-time stationary speech analysis. The frames between adjacent speech frames are shifted by half to one-quarter of the fixed frame length, ensuring that overlapping sampling points are retained between adjacent speech frames and continuously covering the wake-up speech segment. The amplitude envelopes of the left and right channels are extracted within each speech frame. The amplitude envelopes are obtained by taking the absolute value of the speech frame samples and then smoothing them through low-pass filtering. The time corresponding to the maximum point of the amplitude envelope is taken as the arrival time of the left and right channel peaks, resulting in a dual-channel arrival time difference sequence. The arrival time of the peak is refined to sub-sampling precision by three-point parabolic interpolation of the neighborhood of the peak point in the envelope; the left and right directions of the sound source are determined according to the positive and negative distribution of the two-channel time-of-arrival sequence, and the arrival time of the left channel peak is subtracted from the arrival time of the right channel peak to obtain the two-channel time-of-arrival sequence; when the two-channel time-of-arrival sequence is positive, the sound source is to the left, and when it is negative, the sound source is to the right; the lateral angle range corresponding to the candidate sector of the sound source is determined according to the maximum and minimum values of the two-channel time-of-arrival sequence, and the azimuth angle corresponding to the maximum and minimum values is used as the sector boundary; when in continuous speech frames When the positive and negative signs of the two-channel arrival time difference sequence alternate, the azimuth range corresponding to the alternating speech frames is incorporated into the sound source candidate sector, and direct body turning is paused. Body turning is resumed after the current visual candidate target is confirmed as the current interactive target. The azimuth range corresponding to the alternating speech frames refers to the range between the azimuth angles calculated from the maximum and minimum values of the two-channel arrival time difference sequence in the continuous speech frames where the positive and negative signs alternate. Resuming body turning means continuing the process of allocating the main body turning angle, head correction angle, and neck compensation angle.
[0031] Face candidate boxes are extracted within the camera image frame region corresponding to the sound source candidate sector. These boxes are generated by a cascaded classifier using multi-scale sliding window detection within a limited search window. The center coordinates of the face candidate boxes are calculated from their top-left and bottom-right corner coordinates to obtain the center point. The x-coordinate and y-coordinate of the center of each face candidate box are then extracted; the center coordinates are the arithmetic mean of the top-left and bottom-right corner coordinates. The lower half of the face candidate box is designated as the mouth region, and the lower half is cropped downwards along the y-coordinate of the face candidate box center. Kana... The Kanade-Lucas-Tomasi optical flow tracking algorithm tracks feature points in the mouth region to obtain the mouth's motion trajectory. The Kanade-Lucas-Tomasi optical flow tracking algorithm, also known as the corner feature-based local optical flow tracking algorithm, iteratively solves the feature point displacement by minimizing the sum of squares of pixel gray-level differences within a neighborhood window. Feature points are selected from the corner points of the mouth region, and feature points that fail to be tracked do not participate in the generation of the mouth's motion trajectory. The visual candidate target trajectory is formed according to the displacement change of the center point of the face candidate box in consecutive image frames.
[0032] In this implementation scheme, this step realizes the orientational connection between the wake-up voice segment and the camera image frame area, converts the acoustic two-channel arrival time difference sequence into a sound source candidate sector, and extracts the face candidate box, mouth movement trajectory and visual candidate target trajectory on the visual side within the sound source candidate sector, forming a preliminary correspondence between the sound source direction and the visual target, reducing the interference of non-interacting personnel, background faces and environmental noise on target localization, and providing a consistent audio-visual target evidence basis for subsequent calculation of user lateral deviation pixels, user longitudinal deviation pixels, user orientation confidence value and generation of visual verification markers.
[0033] Specifically, the process of obtaining the user's location confidence value through evidence fusion is as follows: Subtract half the image width from the horizontal coordinate of the face candidate box center to obtain the user's horizontal deviation pixels; a positive value indicates the user is on the right side of the image, and a negative value indicates the user is on the left side. Subtract half the image height from the vertical coordinate of the face candidate box center to obtain the user's vertical deviation pixels; a positive value indicates the user is at the bottom of the image, and a negative value indicates the user is at the top. Calculate the proportion of frames in which the face candidate box center is located within the camera image frame region corresponding to the sound source candidate sector out of the total number of frames in which the face candidate box appears, to obtain the sound source candidate box confidence value. The source continuity score is obtained by mapping the candidate sound source sectors to horizontal pixel intervals in the image plane based on the horizontal angle range corresponding to the candidate sound source sectors, the current head yaw angle, and the current body rotation angle. The mouth movement synchronization score is obtained by calculating the proportion of the overlap between the image frame time window corresponding to the peak of the mouth movement trajectory and the wake-up speech segment, where the image frame time window is determined by the image sampling timestamp of the image frame containing the peak and the sampling timestamps of adjacent images. The trajectory continuity score is obtained by calculating the proportion of consecutive frames of the visual candidate target trajectory to the total number of image frames within the current interaction time window. The interaction time window originates from the same interaction timeline. The window start time is the start time of the wake-up voice segment, and the window end time is the image sampling timestamp corresponding to the last frame of the visual candidate target trajectory. Continuous existence means that there are no breaks in the visual candidate target trajectory frames and the re-identification features are consistent. The sound source reception score, mouth movement synchronization score, and trajectory continuity score are fused using the Dempster-Shafer evidence discount fusion method to obtain the user's location confidence value. The sound source reception score, mouth movement synchronization score, and trajectory continuity score are used as the basic probability allocation for the corresponding evidence to support the current interaction target. The remaining probability is allocated to the uncertainty set, and the corresponding evidence score is used as a discount factor in the conflict normalization synthesis. The evidence discount fusion method weakens the conflict evidence according to the discount factor and then merges them using the Dempster-Shafer synthesis rule. When the user's location confidence value is less than the user's location confidence threshold, a visual verification mark is generated, indicating that the current reception target may be misfollowed or the sound source may be unrelated. When the user's location confidence value is greater than or equal to the user's location confidence threshold, the current visual candidate target is confirmed as the current interaction target, which is the desktop companion interaction object locked by the system.
[0034] In this implementation scheme, this step realizes the deviation quantification of the current visual candidate target, audio-visual consistency verification, and interactive target credibility determination. The positional deviation of the face candidate box in the image is converted into user horizontal deviation pixels and user vertical deviation pixels. The sound source reception score, mouth movement synchronization score, and trajectory continuity score are fused into the user orientation credibility value, forming a comprehensive judgment basis that takes into account the sound source direction, vocalization action, and target persistence. This reduces the impact of background faces, non-vocal personnel, short-term false detections, and target chain breaks on interactive target confirmation. It provides a credible target basis for subsequent generation of visual verification markers, construction of head, neck, and body orientation angle constraint sets, and allocation of head yaw reception angle, neck pitch reception angle, and body main steering angle.
[0035] Specifically, the process of constructing the head, neck, and body orientation angle constraint set to obtain the orientation deviation to be eliminated is as follows: Read the user's lateral deviation pixels, longitudinal deviation pixels, user orientation confidence value, and visual verification markers; determine the left and right deviation directions based on the sign of the user's lateral deviation pixels, and determine the pitch deviation direction based on the sign of the user's longitudinal deviation pixels; use the current sampling angle in the angle sequence as the attitude reference, taking the latest angle encoder reading at the orientation allocation start time; use the corresponding limit angle data as the motion boundary, and trim the head yaw angle adjustment range, neck pitch angle adjustment range, and body rotation angle adjustment range according to the deviation direction. The trimming operation will reduce the limit angle data... The difference between the current sampling angle and the current sampling angle is used as the remaining rotation angle in the corresponding direction. The image center region is extracted from the camera image frame sequence. The range in which the center of the face candidate box falls into the image center region is used to determine the camera's field of view coverage area. The display screen's forward orientation angle range corresponding to the body rotation angle is used to determine the inclusion relationship between the display screen's forward orientation angle range and the user's lateral angle range. This forms a set of head, neck, and body orientation angle constraints. The user's lateral angle range is calculated by back-calculating the user's lateral deviation pixels. The inclusion relationship means that the user's lateral angle range falls within the angle range corresponding to the display screen's forward orientation. The set of head, neck, and body orientation angle constraints is a set of joint angle spaces that satisfy both field of view and display constraints.
[0036] Based on the timestamp of the motor control command, single-degree-of-freedom motion segments with only a single motor action and other angle changes within the angle sampling fluctuation range are selected. The angle sampling fluctuation range originates from the angle sequence fluctuation when the corresponding motor has not issued a control command and the motor feedback current sequence is in the static current segment. The difference between the maximum and minimum values of the angle sequence within this time period is taken as the angle sampling fluctuation range. The center displacement of the face candidate box before and after the single-degree-of-freedom motion segment, as well as the difference before and after the corresponding head yaw angle, neck pitch angle, or torso rotation angle, are calculated to form the pixel angle response relationship of the head, neck, and torso. The pixel angle response relationship is a linear fitting relationship with the difference before and after the action as input and the center displacement of the face candidate box as output, and the corresponding slope is obtained through linear least squares fitting. Then, based on the pixel angle response relationship, the pixel angles of the user's lateral deviation pixels and user's longitudinal deviation pixels are inversely calculated to obtain the orientation deviation to be eliminated. The pixel angle inverse calculation divides the pixel deviation by the slope of the pixel angle response relationship to obtain the angle deviation.
[0037] In this implementation plan, this step achieves spatial matching between the head, neck, and torso movement capabilities of the desktop companion robot and the user's deviation requirements. It maps the user's lateral deviation pixels, user's longitudinal deviation pixels, current posture state, mechanism limit boundaries, camera field of view acceptance range, and display screen forward display range into a unified set of head, neck, and body orientation angle constraints, clarifying the orientation adjustment range that each degree of freedom can undertake in the current interaction state. At the same time, it establishes pixel angle response relationships, converting the user's position deviation in the image plane into orientation deviation to be eliminated, providing a calculable, constrainable, and executable allocation basis for the subsequent generation of temporary acceptance sequences, orientation execution sequences, line of sight correction sequences, or limit avoidance marks.
[0038] Specifically, the process of generating temporary receiving sequences, orientation execution sequences, line-of-sight correction sequences, or limit avoidance markers based on the head, neck, and body orientation angle constraint set and the orientation deviation to be eliminated is as follows: Within the action execution window, the time difference from the motor control command timestamp to entering the drive current segment, and the time difference from the drive current segment back to the stationary current segment are calculated based on the motor feedback current sequence. The two types of time differences are added together to obtain the head response duration, neck response duration, and body response duration. The action execution window starts from the motor control command timestamp and ends when the corresponding motor feedback current sequence falls back and remains in the stationary current segment. The window length is determined by the actual execution time of a single action. The drive current segment is the interval in which the feedback current of the motor is continuously higher than the static current baseline after receiving the control command. The portion of the feedback current in the drive current segment that exceeds the static current baseline is greater than the difference between the maximum and minimum values of the feedback current sequence when the motor is not rotating. The static current segment is the interval in which the feedback current fluctuates around the static current baseline when the motor is not rotating. The fluctuation of the feedback current in the static current segment does not exceed the difference between the maximum and minimum values of the feedback current sequence when the motor is not rotating. The static current baseline is determined by the average value of the feedback current when the motor is not rotating.
[0039] Based on the head-neck-body orientation angle constraint set and the orientation deviation to be eliminated, a hierarchical model predictive control process with limit constraints is established. The head-neck-body orientation angle constraint set serves as the limit constraint, and the head yaw acceptance angle, neck pitch acceptance angle, body main steering angle, head correction angle, and neck compensation angle serve as the allocation objects. The first layer aims to recapture the face candidate box by the camera, compares the head response time and neck response time, and determines whether the head yaw angle adjustment range and neck pitch angle adjustment range cover the corresponding orientation deviation to be eliminated. The head yaw angle adjustment range covers the corresponding orientation deviation to be eliminated at the turning angle. In the first layer, the mechanism corresponding to the smaller value between the head response time and the neck response time is selected to perform the field of view reception. The second layer aims to face the user with the display screen and selects the body rotation direction with the smaller absolute value of the rotation angle between the clockwise and counterclockwise directions within the body rotation angle adjustment range. A smaller absolute value of the rotation angle means that the display screen orientation is completed with a shorter path. The third layer aims to reduce redundant compensation and ensures that the head correction angle only compensates for the remaining lateral deviation components not covered by the main body steering angle, ensuring that the lateral deviation ranges addressed by the main body steering angle and the head correction angle are mutually exclusive.
[0040] When visual verification markers exist, head yaw acceptance angle and neck pitch acceptance angle are generated, and a temporary acceptance sequence is output. Specifically, the head yaw acceptance angle is generated from the portion of the user's lateral deviation pixels that can be covered within the head yaw angle adjustment range, and the neck pitch acceptance angle is generated from the portion of the user's vertical deviation pixels that can be covered within the neck pitch angle adjustment range. When no visual verification markers exist and the forward display area of the screen does not cover the user's lateral angle range calculated from the user's lateral deviation pixels, the body main turning angle, head correction angle, and neck compensation angle are generated. Specifically, the body main turning angle is generated from the deviation between the forward display area of the screen and the center of the face candidate box corresponding to the current interactive target, and the head correction angle is generated from the deviation between the forward display area of the screen and the center of the face candidate box corresponding to the current interactive target. The remaining lateral and longitudinal deviations generate the head correction angle and neck compensation angle, and output the orientation execution sequence. When there are no visual verification marks and the display screen's forward display area covers the user's lateral angle range calculated from the user's lateral deviation pixels, the remaining lateral and longitudinal deviations generate the head correction angle and neck compensation angle respectively, and output the line of sight correction sequence. When the head yaw acceptance angle or head correction angle exceeds the head yaw angle adjustment range, the neck pitch acceptance angle or neck compensation angle exceeds the neck pitch angle adjustment range, or the body main steering angle exceeds the body rotation angle adjustment range, a limit avoidance mark is output. The limit avoidance mark indicates that the current angle to be executed exceeds the physical limit of the mechanism and an alternative strategy needs to be called.
[0041] In this embodiment, Table 1 is a multi-degree-of-freedom orientation allocation result data table, recording the user's lateral deviation pixels, user's longitudinal deviation pixels, user's orientation confidence value, visual verification mark, allocation judgment result, and output result for five interaction samples. The allocation judgment result indicates the allocation status of the current sample after entering temporary verification, display screen coverage judgment, or reachability interval judgment. For sample one, the user's lateral deviation pixels are 0.62, the user's longitudinal deviation pixels are 0.18, the user's orientation confidence value is 0.44, the visual verification mark is present, the allocation judgment result is "requires verification," and the output result is a temporary acceptance sequence. For sample two, the user's lateral deviation pixels are 0.55, the user's longitudinal deviation pixels are 0.20, the user's orientation confidence value is 0.78, the visual verification mark is absent, the allocation judgment result is "display screen not covered," and the output result is an orientation execution sequence. For sample three, the user's lateral deviation pixels are 0.12, the user's longitudinal deviation pixels are -0.16, and the user's orientation confidence value is 0.8. 4. The visual verification mark is absent, the allocation judgment result is that the display screen is covered, and the output result is a line-of-sight correction sequence; Sample 4 has a user lateral deviation pixel of 0.91, a user longitudinal deviation pixel of 0.34, and a user orientation confidence value of 0.79. The visual verification mark is absent, the allocation judgment result is that the reachable range is insufficient, and the output result is a limit avoidance mark; Sample 5 has a user lateral deviation pixel of -0.48, a user longitudinal deviation pixel of 0.11, and a user orientation confidence value of 0.82. The visual verification mark is absent, the allocation judgment result is that the display screen is not covered, and the output result is an orientation execution sequence. As can be seen from Table 1, different user deviations, user orientation confidence values, and allocation judgment results will correspond to different orientation allocation outputs.
[0042] Table 1. Multi-DOF Orientation Assignment Results
[0043] like Figure 3 The figure shows a scatter plot of user deviation pixels and orientation assignment results. The horizontal axis represents the user's lateral deviation pixels, and the vertical axis represents the user's vertical deviation pixels. Each scatter point corresponds to an interaction sample in Table 1. In the figure, square markers represent temporary acceptance sequences, circular markers represent orientation execution sequences, triangle markers represent line-of-sight correction sequences, and diamond markers represent limit avoidance markers. The horizontal and vertical dashed lines represent the zero-value baselines for user vertical deviation pixels and user lateral deviation pixels, respectively. As can be seen from Table 1, although Sample 1 has lateral and vertical deviations, it outputs a temporary acceptance sequence because of the presence of visual verification markers. Sample 4 is located in the area with large deviations and outputs limit avoidance markers. Samples 2 and 5 are located in the left and right deviation areas, respectively, and both output orientation execution sequences.
[0044] like Figure 4The figure shows the distribution of user orientation confidence values and lateral deviations. The horizontal axis represents the user's lateral deviation in pixels, and the vertical axis represents the user's orientation confidence value. Each scatter point corresponds to an interaction sample in Table 1. The black dotted horizontal lines in the figure represent the user orientation confidence threshold, used to distinguish whether a visual verification mark is generated; the gray dashed lines are the sample distribution contour lines, used to represent the density of sample distribution on the user's lateral deviation pixel and user orientation confidence value planes, and the gray filled areas are the distribution background corresponding to the contour lines. Referring to Table 1, it can be seen that the user orientation confidence value of sample one is lower than the user orientation confidence threshold, and a temporary acceptance sequence is output; the user orientation confidence values of samples two through five are all higher than the user orientation confidence threshold, and the visual verification mark does not exist. The corresponding result is then output based on the display screen coverage status or reachability interval status.
[0045] In this implementation plan, this step achieves coordinated allocation of response speed, reachability, display orientation, and camera field of view among the head, neck, and torso. It transforms the orientation deviation to be eliminated into specific execution results that conform to the mechanism's limits and the interactive target state, avoiding repeated compensation or over-limit execution of the same deviation by the head, neck, and torso. When the target credibility is insufficient, a temporary acceptance sequence is output; when the display is not facing the user, an orientation execution sequence is output; when the display has covered the user's direction, a line-of-sight correction sequence is output; and when the mechanism's reachability is insufficient, a limit avoidance mark is output. This ensures that the head yaw acceptance angle, neck pitch acceptance angle, torso main steering angle, head correction angle, and neck compensation angle have clear allocation criteria and execution boundaries, improving the stability, response efficiency, and human-computer interaction continuity of the desktop companion robot's orientation control.
[0046] Specifically, the process of generating motor control commands based on temporary acceptance sequences, orientation execution sequences, or line-of-sight correction sequences is as follows: (e.g.) Figure 5 The diagram shows the multi-degree-of-freedom motor command conversion and limit avoidance execution flowchart. According to the temporary acceptance sequence, orientation execution sequence or line of sight correction sequence, the yaw angle assigned to the head, the pitch angle assigned to the neck and the rotation angle assigned to the torso are converted into head motor control commands, neck motor control commands and torso motor control commands respectively, and sent to the motor drive board. The conversion from angle to motor command is completed according to the angle control format preset by the motor drive board.
[0047] When a limit avoidance marker exists, a replacement acceptance angle is generated based on the remaining angles of the mechanisms that do not exceed the head yaw angle adjustment range, neck pitch angle adjustment range, or body rotation angle adjustment range. The head and body are used to replace the execution angles that exceed the range corresponding to the lateral deviation, and the neck is used to replace the execution angles that exceed the range corresponding to the longitudinal deviation. When there are multiple mechanisms that do not exceed the range for the same deviation, a replacement mechanism is determined according to the allocation order in the temporary acceptance sequence, orientation execution sequence, or line-of-sight correction sequence, and the execution angles that exceed the range are replaced with the replacement acceptance angles to bring the face candidate box into the camera's field of view acceptance range. The remaining angle is the difference between the corresponding motion boundary and the current sampling angle. When there is no limit avoidance marker, the corresponding mechanism actions are executed in the order of the temporary acceptance sequence, orientation execution sequence, or line-of-sight correction sequence. Each mechanism action is triggered serially in the sequence and waits for the previous action to be completed and confirmed before starting the next action.
[0048] In this implementation plan, the orientation assignment result is transformed into the motor execution action. The temporary acceptance sequence, orientation execution sequence and gaze correction sequence are transformed into control actions that can be executed by the head, neck and torso. Limit avoidance is triggered when the mechanism approaches or exceeds the movement boundary, ensuring that the face candidate box re-enters the camera's field of view acceptance range, avoiding excessive rotation, action conflict and invalid compensation, and improving the execution safety, action continuity and target acceptance stability of multi-degree-of-freedom orientation control.
[0049] Specifically, the process of executing orientation control and outputting the orientation completion mark is as follows: During execution, the angle sequence, motor feedback current sequence, and camera image frame sequence are continuously read; the target angle in the motor control command is compared with the final sampled angle of the angle sequence after the corresponding control command is executed, and the absolute difference is calculated to obtain the head execution deviation, neck execution deviation, and body execution deviation, respectively. The user orientation confidence value is then recalculated based on the camera image frame sequence after execution, the sound source candidate sector, the mouth movement trajectory, and the visual candidate target trajectory, according to the fusion process of sound source reception, mouth movement synchronization, and trajectory continuity.
[0050] Orientation control process: When the updated user orientation confidence value is greater than or equal to the user orientation confidence threshold, maintain the current head yaw angle, neck pitch angle, and body rotation angle; when the updated user orientation confidence value is less than the user orientation confidence threshold, sequentially execute head left boundary pause frame acquisition, head middle position pause frame acquisition, and head right boundary pause frame acquisition within the sound source candidate sector. The head left boundary and head right boundary are the two boundaries after the intersection of the corresponding lateral angle range of the sound source candidate sector and the head yaw angle adjustment range. The head middle position is the middle angle between the two boundaries. Pause frame acquisition refers to pausing the movement after controlling the head to reach the target angle and continuously acquiring several frames for orientation reconfirmation; if the recalculated user orientation confidence value is still less than the user orientation confidence threshold, then execute neck downward... The system captures frames for both tilting and neck-tilt pauses, as well as tilting and neck-up pauses, corresponding to the tilt and neck-up limits of the neck pitch angle adjustment range, respectively. If the lateral angle range corresponding to the candidate sector of the sound source exceeds the head yaw angle adjustment range, the body is controlled to perform segmented compensation rotation along the direction of the candidate sector of the sound source, and the current interaction target is reconfirmed. The segmented compensation rotation divides the angle difference exceeding the head yaw angle adjustment range into multiple body compensation rotation angles, and verifies each segment by capturing frames. After each segment of compensation rotation, the user's orientation confidence value is recalculated, and the current interaction target is reconfirmed based on the user's orientation confidence value being greater than or equal to the user's orientation confidence threshold. After the orientation control process is completed, the system outputs an orientation completion marker, orientation deviation record, degree of freedom limit record, and interaction acceptance status.
[0051] In this implementation plan, this step achieves real-time feedback verification and secondary acceptance correction of the orientation execution process. It incorporates head execution deviation, neck execution deviation, body execution deviation, and the updated user orientation confidence value into the closed-loop judgment, promptly identifying situations such as motor incomplete execution, target deviation, and audio-visual acceptance failure. When the user orientation confidence value meets the requirements, the current posture is maintained. When the user orientation confidence value is insufficient, head pause frame acquisition, neck pause frame acquisition, and body segment compensation rotation are sequentially invoked to reconfirm the current interaction target. This enables the desktop companion robot to maintain stable alignment of the camera acquisition direction, display screen display direction, and voice interaction direction during continuous interaction, improving the closed-loop reliability of orientation control, target retention capability, and interaction continuity.
[0052] like Figure 2As shown, the second aspect of the present invention provides a desktop companion robot's audio-visual composite positioning and orientation control system, comprising: an acquisition and preprocessing unit, used to acquire raw data across sensors and perform time alignment, data correction, anomaly removal, and standardization to form an audio-visual attitude synchronization data chain; an audio-visual composite orientation receiving unit, used to generate candidate sectors of sound sources and extract visual target features based on the audio-visual attitude synchronization data chain, and obtain user orientation confidence values through evidence fusion; a multi-degree-of-freedom orientation allocation unit, used to construct a head, neck, and body orientation angle constraint set, obtain the orientation deviation to be eliminated, and generate a temporary receiving sequence, orientation execution sequence, line-of-sight correction sequence, or limit avoidance mark based on the head, neck, and body orientation angle constraint set and the orientation deviation to be eliminated; and a segmented execution and closed-loop correction unit, used to generate motor control commands based on the temporary receiving sequence, orientation execution sequence, or line-of-sight correction sequence, execute orientation control, and output an orientation completion mark.
[0053] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0054] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A method for acoustic-visual composite positioning and orientation control of a desktop companion robot, characterized in that, Includes the following steps: S1 collects raw data across sensors and performs time alignment, data correction, anomaly removal, and standardization to form an audio-visual attitude synchronization data chain. S2, based on the audio-visual attitude synchronization data link, generate candidate sectors of sound sources and extract visual target features, and obtain the user's location confidence value through evidence fusion; S3, construct the head, neck and body orientation angle constraint set, obtain the orientation deviation to be eliminated, and generate temporary acceptance sequence, orientation execution sequence, line of sight correction sequence or limit avoidance mark based on the head, neck and body orientation angle constraint set and the orientation deviation to be eliminated; S4 generates motor control commands based on the temporary acceptance sequence, orientation execution sequence, or line-of-sight correction sequence, executes orientation control, and outputs an orientation completion marker.
2. The method for acoustic-visual composite positioning and orientation control of a desktop companion robot according to claim 1, characterized in that: The specific process for collecting raw data across sensors is as follows: The data collection target is a desktop companion robot equipped with a head display screen, head camera, left microphone, right microphone, neck pitch mechanism, body rotation mechanism and motor drive board; The system collects raw data across multiple sensors, including: simultaneous acquisition of left and right microphone speech sampling sequences and acoustic sampling timestamps via the left and right microphones; acquisition of camera image frame sequences, image sampling timestamps, image width, and image height via the head camera; acquisition of angle sequences via head, neck, and torso angle encoders, including head yaw angle sequences, neck pitch angle sequences, and torso rotation angle sequences; acquisition of limit angle data via limit switches and angle encoders during robot self-check actions; and acquisition of motor control command timestamps and motor feedback current sequences via the motor drive board, including head motor feedback current sequences, neck motor feedback current sequences, and torso motor feedback current sequences.
3. The method for acoustic-visual composite positioning and orientation control of a desktop companion robot according to claim 1, characterized in that: The specific process of performing time alignment, data correction, anomaly removal, and standardization to form an audio-visual attitude synchronization data link is as follows: Based on the acoustic sampling timestamp and image sampling timestamp, the left microphone speech sampling sequence, the right microphone speech sampling sequence, and the camera image frame sequence are arranged on the same interactive time axis. Silent segment clipping, clipped segment removal, and left and right channel amplitude correction are performed on the left and right microphone speech sampling sequences. Jitter frame removal, exposure consistency, and distortion correction are performed on the camera image frame sequence. Jump points are removed from the angle sequence and the motor feedback current sequence to form an audio-visual attitude synchronization data link. The left microphone speech sampling sequence, the right microphone speech sampling sequence, and the motor feedback current sequence are standardized using the Z-score normalization algorithm. The pixel brightness values and angle sequences in the camera image frame sequence are normalized using the Min-Max normalization algorithm.
4. The method for acoustic-visual composite positioning and orientation control of a desktop companion robot according to claim 1, characterized in that: The specific process of generating candidate sectors for sound sources and extracting visual target features based on the audio-visual attitude synchronization data link is as follows: Extract wake-up speech segments from the audio-visual posture synchronization data link; divide the left microphone speech sampling sequence and the right microphone speech sampling sequence into multiple speech frames with a fixed frame length; extract the amplitude envelopes of the left and right channels in each speech frame; The time corresponding to the point of maximum amplitude envelope is taken as the arrival time of the left channel peak and the arrival time of the right channel peak. The arrival time difference sequence of the two channels is obtained by subtracting the arrival time of the left channel peak from the arrival time of the right channel peak. The left and right directions of the sound source are determined based on the positive and negative distribution of the two-channel time difference of arrival sequence; the lateral angle range corresponding to the candidate sector of the sound source is determined based on the maximum and minimum values of the two-channel time difference of arrival sequence; when the positive and negative signs of the two-channel time difference of arrival sequence alternate in continuous speech frames, the azimuth angle range corresponding to the alternating speech frames is incorporated into the candidate sector of the sound source, and direct body turning is paused. Body turning is resumed after the current visual candidate target is confirmed as the current interactive target. Extract face candidate boxes within the camera image frame region corresponding to the sound source candidate sector. Calculate the center coordinates from the coordinates of the upper left and lower right corners of the face candidate boxes to obtain the center point of the face candidate boxes. Then extract the horizontal and vertical coordinates of the center of the face candidate boxes. The lower half of the face candidate box is taken as the mouth region. The Kanade-Lucas-Tomasi optical flow tracking algorithm is used to track the feature points of the mouth region to obtain the mouth motion trajectory. The visual candidate target trajectory is formed according to the displacement change of the center point of the face candidate box in consecutive image frames.
5. The method for acoustic-visual composite positioning and orientation control of a desktop companion robot according to claim 1, characterized in that: The specific process of obtaining the user's location confidence value through evidence fusion is as follows: The horizontal deviation pixels of the user are obtained by subtracting half the image width from the horizontal coordinate of the face candidate box center; the vertical deviation pixels of the user are obtained by subtracting half the image height from the vertical coordinate of the face candidate box center; the sound source continuity score is obtained by calculating the proportion of frames in which the center of the face candidate box is located within the camera image frame region corresponding to the sound source candidate sector; the mouth movement synchronization score is obtained by calculating the proportion of the overlap duration between the image frame time window corresponding to the peak of the mouth movement trajectory and the wake-up speech segment; the trajectory continuity score is obtained by calculating the proportion of the number of frames in which the visual candidate target trajectory exists continuously to the total number of image frames in the current interaction time window; and the user's location confidence value is obtained by fusing the sound source continuity score, mouth movement synchronization score, and trajectory continuity score using the Dempster-Shafer evidence discount fusion method. When the user's location confidence value is less than the user's location confidence threshold, a visual verification mark is generated. When the user's location confidence value is greater than or equal to the user's location confidence threshold, the current visual candidate target is confirmed as the current interaction target.
6. The method for acoustic-visual composite positioning and orientation control of a desktop companion robot according to claim 1, characterized in that: The specific process of constructing the head-neck-body orientation angle constraint set to obtain the orientation deviation to be eliminated is as follows: Read the user's lateral deviation pixels, user's longitudinal deviation pixels, user's orientation confidence value, and visual verification markers. Determine the left and right deviation directions based on the sign of the user's lateral deviation pixels, and determine the pitch deviation direction based on the sign of the user's longitudinal deviation pixels. Use the current sampling angle in the angle sequence as the attitude reference. Use the corresponding limit angle data as the motion boundary, and trim the head yaw angle adjustment range, neck pitch angle adjustment range, and body rotation angle adjustment range according to the deviation direction. Extract the image center region from the camera image frame sequence. Determine the camera's field of view coverage area based on the range where the center of the face candidate box falls into the image center region. Determine the display screen's forward display area based on the inclusion relationship between the display screen's front facing angle range corresponding to the body rotation angle and the user's lateral angle range, forming a set of head, neck, and body facing angle constraints. Based on the timestamp of the motor control command, single-degree-of-freedom motion segments with only a single motor action and other angle changes within the angle sampling fluctuation range are selected. The displacement of the face candidate box center before and after a single degree of freedom action segment is calculated, as well as the difference in the head yaw angle, neck pitch angle or torso rotation angle before and after the action, to form the pixel angle response relationship of the head, neck and torso; then, based on the pixel angle response relationship, the pixel angle of the user's lateral deviation pixel and the user's longitudinal deviation pixel is calculated to obtain the orientation deviation to be eliminated.
7. The method for acoustic-visual composite positioning and orientation control of a desktop companion robot according to claim 1, characterized in that: The specific process of generating temporary acceptance sequences, orientation execution sequences, line-of-sight correction sequences, or limit avoidance markers based on the head, neck, and body orientation angle constraint set and the orientation deviation to be eliminated is as follows: Within the action execution window, the time difference from the motor control command timestamp to entering the drive current segment and the time difference from the drive current segment back to the stationary current segment are calculated based on the motor feedback current sequence. The two types of time differences are added together to obtain the head response time, neck response time and body response time. Based on the head, neck and body orientation angle constraint set and the orientation deviation to be eliminated, a hierarchical model prediction control process with limit constraints is established: The first layer takes the camera recaptures the face candidate box as the goal, compares the head response time and the neck response time, and determines whether the head yaw angle adjustment range and the neck pitch angle adjustment range cover the corresponding orientation deviation to be eliminated. Among the heads or necks that cover the corresponding orientation deviation to be eliminated, the mechanism corresponding to the smaller value of the head response time and the neck response time is selected to perform the field of view transfer. The second layer aims to make the display screen face the user and selects the body rotation direction with the smaller absolute value of the rotation angle between the clockwise and counterclockwise directions within the body rotation angle adjustment range; the third layer aims to reduce redundant compensation and makes the head correction angle only compensate for the remaining lateral deviation components not covered by the main body steering angle. When a visual verification mark exists, generate the head yaw acceptance angle and neck pitch acceptance angle, and output a temporary acceptance sequence; when no visual verification mark exists and the display screen's forward display area does not cover the user's lateral angle range calculated from the user's lateral deviation pixels, generate the body main steering angle, head correction angle, and neck compensation angle, and output an orientation execution sequence; when no visual verification mark exists and the display screen's forward display area covers the user's lateral angle range calculated from the user's lateral deviation pixels, generate the head correction angle and neck compensation angle, and output a line-of-sight correction sequence; when the head yaw acceptance angle or head correction angle exceeds the head yaw angle adjustment range, the neck pitch acceptance angle or neck compensation angle exceeds the neck pitch angle adjustment range, or the body main steering angle exceeds the body rotation angle adjustment range, output a limit avoidance mark.
8. The method for acoustic-visual composite positioning and orientation control of a desktop companion robot according to claim 1, characterized in that: The specific process for generating motor control commands based on temporary acceptance sequences, orientation execution sequences, or line-of-sight correction sequences is as follows: According to the temporary takeover sequence, orientation execution sequence, or line-of-sight correction sequence, the yaw angle assigned to the head, the pitch angle assigned to the neck, and the rotation angle assigned to the torso are converted into head motor control commands, neck motor control commands, and torso motor control commands, respectively, and then sent to the motor drive board. When there is a limit avoidance mark, a substitute acceptance angle is generated based on the remaining angle of the mechanism that does not exceed the head yaw angle adjustment range, neck pitch angle adjustment range, or body rotation angle adjustment range. The execution angle that exceeds the range is replaced by the substitute acceptance angle, so that the face candidate box enters the camera's field of view acceptance range. When there is no limit avoidance mark, the corresponding mechanism action is executed in the order of temporary acceptance sequence, orientation execution sequence, or line of sight correction sequence.
9. The method for acoustic-visual composite positioning and orientation control of a desktop companion robot according to claim 1, characterized in that: The specific process of performing orientation control and outputting an orientation completion marker is as follows: During execution, the angle sequence, motor feedback current sequence, and camera image frame sequence are read. The target angle in the motor control command is compared with the final sampled angle of the angle sequence after the corresponding control command is executed. The absolute difference is calculated to obtain the head execution deviation, neck execution deviation, and body execution deviation, and the user orientation confidence value is recalculated. Orientation control process: When the updated user orientation confidence value is greater than or equal to the user orientation confidence threshold, maintain the current head yaw angle, neck pitch angle, and body rotation angle; when the updated user orientation confidence value is less than the user orientation confidence threshold, sequentially execute head left boundary stay frame sampling, head middle position stay frame sampling, and head right boundary stay frame sampling within the sound source candidate sector. If the recalculated user orientation confidence value is still less than the user orientation confidence threshold, then perform neck downward pause frame acquisition and neck upward pause frame acquisition. If the lateral angle range corresponding to the candidate sector of the sound source exceeds the head yaw angle adjustment range, the body is controlled to rotate in segments along the direction of the candidate sector of the sound source, and the current interactive target is reconfirmed. After the orientation control process is completed, the output includes an orientation completion marker, orientation deviation record, degree of freedom limit record, and interactive connection status.
10. A desktop companion robot's acoustic-visual composite positioning and orientation control system, applied to the desktop companion robot's acoustic-visual composite positioning and orientation control method according to any one of claims 1-9, characterized in that, include: The acquisition and preprocessing unit is used to acquire raw data across sensors and perform time alignment, data correction, anomaly removal and standardization to form an audio-visual attitude synchronization data chain. The audio-visual composite orientation receiving unit is used to generate candidate sectors of sound sources and extract visual target features based on the audio-visual attitude synchronization data link, and obtain the user's orientation confidence value through evidence fusion. The multi-degree-of-freedom orientation assignment unit is used to construct a set of head, neck and body orientation angle constraints, obtain the orientation deviation to be eliminated, and generate temporary acceptance sequences, orientation execution sequences, line of sight correction sequences or limit avoidance marks based on the set of head, neck and body orientation angle constraints and the orientation deviation to be eliminated. The segmented execution and closed-loop correction unit is used to generate motor control commands based on temporary acceptance sequences, orientation execution sequences, or line-of-sight correction sequences, execute orientation control, and output an orientation completion marker.
Citation Information
Patent Citations
A robot system and method based on intelligent sound source localization and voice control
CN107199572B
A robot control method, storage medium, and robot that can be used on demand
CN111055288B