Intelligent interaction method and device of cleaning robot, medium and electronic equipment

By having a cleaning robot actively accompany the user during video communication and adjust its position and the orientation of the image acquisition module, the problem of video interruption caused by user movement was solved, thus achieving continuity and stability in video communication.

CN121370007BActive Publication Date: 2026-05-01DREAM INNOVATION TECH (SUZHOU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DREAM INNOVATION TECH (SUZHOU) CO LTD
Filing Date
2025-12-25
Publication Date
2026-05-01

Smart Images

  • Figure CN121370007B_ABST
    Figure CN121370007B_ABST
Patent Text Reader

Abstract

The application provides an intelligent interaction method and device of a cleaning robot, a medium and an electronic device. The method comprises: in response to a video communication starting instruction, establishing a connection with a communication opposite end; in an environment, identifying a local communication object of the video communication based on a pre-stored feature; and when performing a video communication task, actively accompanying the local communication object at a preset accompanying distance and making the local communication object located in a video picture transmitted to the communication opposite end. The cleaning robot actively accompanies the local communication object during the video call, thereby solving the problem of video picture loss or interruption caused by the movement of the local communication object and making the mobile video communication continuous and stable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more specifically, to an intelligent interaction method, device, medium, and electronic device for a cleaning robot. Background Technology

[0002] With the trend of multi-functional integration in cleaning robots, some products have already incorporated autonomous cleaning and video communication functions.

[0003] In related technologies, when cleaning robots implement video communication functions, they are in a static position and cannot dynamically and actively accompany the user (the communication target) as the user moves. This results in the interruption of the user's view on the other end of the communication once the user leaves the fixed initial shooting area.

[0004] Therefore, it is evident that enabling cleaning robots to proactively and dynamically accompany users in real-time during video communication has become a pressing technical issue that needs to be addressed to improve service reliability and user experience. Summary of the Invention

[0005] The purpose of this application is to provide an intelligent interaction method, device, medium, and electronic device for a cleaning robot, enabling a dynamic, real-time, and proactive intelligent interaction scheme for accompanying users during video communication, thereby ensuring that the communication peer can continuously and stably acquire the user's video feed. The specific solution is as follows:

[0006] According to a specific embodiment of this application, in a first aspect, this application provides an intelligent interaction method for a cleaning robot, the method comprising:

[0007] In response to a video communication initiation command, establish a connection with the communication peer;

[0008] Based on pre-stored features in the environment, identify the local communication object in video communication;

[0009] When performing a video communication task, the local communication object is actively accompanied at a preset accompanying distance, and the local communication object is placed in the video frame transmitted to the communication peer.

[0010] In some possible embodiments, the step of actively accompanying the local communication object at a preset accompanying distance and placing the local communication object in the video frame transmitted to the communication peer includes:

[0011] Based on the video stream of video communication, adjust the position of the cleaning robot and / or the orientation of the image acquisition module of the cleaning robot to maintain the frontal image of the local communication object in a preset position on the display screen of the communication counterpart.

[0012] In some possible embodiments, adjusting the position of the cleaning robot and / or the orientation of the image acquisition module of the cleaning robot includes:

[0013] Based on the video stream from video communication, the facial orientation of the local communication object is obtained;

[0014] Based on the deviation between the facial orientation and the facing direction, the position of the cleaning robot and / or the orientation of the image acquisition module are adjusted so that the frontal image of the local communication object is located at a preset display position on the other end of the communication.

[0015] In some possible embodiments, adjusting the position of the cleaning robot and / or the orientation of the image acquisition module of the cleaning robot includes:

[0016] Identify the current position of the face of the local communication object in the video frame;

[0017] Calculate the offset between the current position and the preset position on the local screen;

[0018] Based on the offset, adjust the position of the cleaning robot and / or the orientation of the image acquisition module so that the frontal image of the local communication object is located at the preset display position of the communication peer.

[0019] In some possible embodiments, adjusting the position of the cleaning robot and / or the orientation of the image acquisition module to position the frontal image of the local communication object at a preset display position on the other end includes:

[0020] When it is determined that the cleaning robot cannot move, the orientation of the image acquisition module is adjusted so that the frontal image of the local communication object is located at a preset display position on the other end of the communication; or,

[0021] When determining the movable position of the cleaning robot, the priority of the orientation of the image acquisition module and the position adjustment of the cleaning robot is determined based on the energy consumption of the cleaning robot's position adjustment and the orientation adjustment of the image acquisition module.

[0022] In some possible embodiments, the step of actively accompanying the local communication object at a preset accompanying distance and placing the local communication object in the video frame transmitted to the communication peer includes:

[0023] When the facial image of the local communication object fails to meet the preset communication quality conditions, an image frame containing a visually enhanced facial image is generated based on the pre-stored facial feature information of the local communication object, and output to the communication peer.

[0024] In some possible embodiments, the failure to meet the preset communication quality conditions includes at least one of the following situations:

[0025] Facial loss;

[0026] The area of ​​the face that is obscured exceeds a preset threshold;

[0027] The face is not shown from a frontal angle;

[0028] The facial image clarity is below the preset threshold.

[0029] In some possible embodiments, generating the image frame containing the visually enhanced facial image includes:

[0030] Generate a fitted facial image that matches the current head posture and body orientation of the local communication object;

[0031] The fitted facial image is fused with the background of the current image frame to generate and output the image frame containing the visually enhanced facial image.

[0032] In some possible embodiments, before fusing the fitted facial image with the background of the current image frame, the method further includes:

[0033] Evaluate the visual consistency coefficient between the fitted facial image and the body region of the local communication object in the current frame;

[0034] Based on the visual coordination, facial fusion or body replacement is performed.

[0035] In some possible embodiments, performing facial fusion or body replacement based on the visual coordination includes:

[0036] When the visual coordination coefficient is higher than a preset threshold, the fitted facial image is fused with the original body region in the current frame;

[0037] When the visual coordination coefficient is lower than or equal to a preset threshold, a pre-stored half-body or full-body reference image containing a frontal face is used to replace the body region of the local communication object in the current frame and is blended with the background of the current image frame.

[0038] In some possible embodiments, the method further includes: during video communication, responding to an active companion behavior mode switching instruction, and switching the current active companion behavior from a first active companion mode to a second active companion mode according to the active companion behavior mode switching instruction;

[0039] Among them, the first active companion mode and the second active companion mode differ in at least one of the active companion parameters, namely, companion distance, active companion speed, camera focal length, and camera pitch angle.

[0040] In some possible embodiments, the active companion parameters of the first active companion mode position the face of the local communication object at the center of the display screen of the communication peer; the active companion parameters of the second active companion mode position the entire body of the local communication object at the center of the display screen of the communication peer.

[0041] In some possible embodiments, the video communication activation command includes at least one of the following:

[0042] Receive video communication commands input by the user through the operation interface of the cleaning robot body;

[0043] The cleaning robot's local voice interaction module recognizes video communication commands issued by the user.

[0044] Receive a video communication establishment request command from the terminal device;

[0045] Upon receiving a call from a preset contact, a video communication initiation command is generated and executed.

[0046] In some possible embodiments, identifying the local communication object of the video communication based on pre-stored features in the environment includes:

[0047] Based on the video communication start command, the identity information of the communication object on this end is determined;

[0048] Based on the identity information, retrieve the pre-stored features corresponding to the identity information from the pre-stored biometric database;

[0049] Based on the matching results between the pre-stored features and the biometric features of humanoid targets detected in the environment, the local communication object of the video communication is identified;

[0050] Based on the matching results, the local communication object of the video communication is identified.

[0051] In some possible embodiments, identifying the local communication target of the video communication based on the matching result of the pre-stored features and the biometric features of humanoid targets detected in the environment includes:

[0052] Detect humanoid targets in the environment and collect multimodal biometrics of each humanoid target;

[0053] The pre-stored features are fused and matched with the multimodal biometric features of each humanoid target, and the local communication object of the video communication is identified based on the matching result.

[0054] In some possible embodiments, the detection of humanoid targets in the environment includes:

[0055] Control the cleaning robot to move within the environment and collect environmental data in real time;

[0056] Based on the human silhouette features contained in the environmental data, human-shaped targets are detected in the environment.

[0057] In some possible embodiments, the multimodal biometrics include at least two of image features, voiceprint features, and gait features; wherein the image features include facial features or human skeletal keypoint features.

[0058] In some possible embodiments, the method further includes:

[0059] During video communication, the cleaning robot responds to ancillary task instructions and executes the ancillary tasks corresponding to those instructions.

[0060] After the auxiliary task is completed, the cleaning robot is controlled to adjust its own position and / or camera orientation so that the frontal image of the local communication object is located at the preset display position of the communication counterpart.

[0061] During the execution of the ancillary tasks, a video communication connection with the communication peer is maintained; the ancillary tasks include at least one of the following:

[0062] Perform local cleaning at the current location of the local communication object;

[0063] Move to a preset location and collect environmental data at that preset location;

[0064] The output module of the cleaning robot outputs preset audio or image content that is independent of the video communication.

[0065] In some possible embodiments, the method further includes: when waiting for or executing a first independent task, and receiving a trigger instruction for a second independent task, determining a scheduling scheme for the first independent task and the second independent task based on a preset task arbitration decision model, according to the identity information of the user associated with the first independent task and the identity information of the initiator of the second independent task, as well as the preset execution priorities of the first independent task and the second independent task;

[0066] The scheduling scheme includes:

[0067] Interrupt the execution of the first independent task and switch to the second independent task;

[0068] Maintain execution of the first independent task, but do not execute the second independent task; or,

[0069] Send a confirmation request to the user associated with the first independent task, and switch to the second independent task or continue executing the first independent task based on the feedback.

[0070] In some possible embodiments, the decision rules of the task arbitration decision model include:

[0071] Based on the preset permission and identity levels of the user associated with the first independent task and the initiator of the second independent task, a scheduling scheme is determined.

[0072] When the permission levels of the identity information of the user associated with the first independent task and the identity information of the initiator of the second independent task are the same, a scheduling scheme is determined based on the preset execution priorities of the first independent task and the second independent task.

[0073] In some possible embodiments, determining the scheduling scheme based on the preset execution priorities of the first independent task and the second independent task includes at least one of the following:

[0074] When the first independent task is a video communication task, output a scheduling scheme to maintain the execution of the first independent task;

[0075] When the preset execution priorities of the first independent task and the second independent task are the same, a confirmation query is sent to the user associated with the first independent task, and the scheduling scheme of switching to the second independent task or maintaining the execution of the first independent task is switched according to the feedback result.

[0076] When the preset execution priorities of the first independent task and the second independent task are different, the scheduling scheme of the task with the higher execution priority is output.

[0077] In some possible embodiments, the first independent task and the second independent task are both any one of a cleaning task, a video communication task, and an interactive task; wherein, the interactive task includes: recording the behavior of an interactive object, and outputting at least one of motion feedback, sound feedback, light effect feedback, and screen feedback based on the behavior.

[0078] In some possible embodiments, the action interaction objects include humans and animals.

[0079] According to a specific embodiment of this application, in a second aspect, this application also provides a cleaning robot, comprising:

[0080] The communication connection unit is configured to establish a connection with the communication peer in response to a video communication start command.

[0081] The target search unit is configured to identify the local communication object of the video communication based on pre-stored features in the environment;

[0082] The companion control unit is configured to actively accompany the local communication object at a preset companion distance when performing a video communication task, and to place the local communication object in the video frame transmitted to the communication peer.

[0083] According to a specific embodiment of this application, in a third aspect, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any of the preceding claims.

[0084] According to a specific embodiment of this application, in a fourth aspect, this application also provides an electronic device, including: one or more processors; and a storage device configured to store one or more programs, which, when executed by the one or more processors, cause the one or more processors to perform the method as described in any of the preceding claims.

[0085] Compared with the prior art, the above-described solutions of this application have at least the following beneficial effects:

[0086] By responding to a video communication initiation command, a connection is established with the communication peer. After the connection is established, the local communication object is identified based on pre-stored features in the environment, thus formally initiating the video communication task between the local communication object and the communication peer. During the execution of the video communication task, the local communication object is actively accompanied at a preset distance, ensuring that the local communication object is within the video frame transmitted to the communication peer. Therefore, the cleaning robot is no longer a fixed monitoring probe, but an intelligent videographer that actively accompanies the user's movement. It can automatically lock onto the video communication object and maintain an appropriate distance for active accompaniment, ensuring that the other party is always clearly visible in the frame. This solves the problem of video frame loss or interruption caused by the user's (local communication object's) movement, making video communication while in motion continuous and stable. Attached Figure Description

[0087] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and are configured together with the description to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:

[0088] Figure 1 A flowchart illustrating an intelligent interaction method for a cleaning robot provided in an embodiment of this application;

[0089] Figure 2 Another flowchart illustrating an intelligent interaction method for a cleaning robot provided in an embodiment of this application;

[0090] Figure 3 A schematic diagram of an intelligent interactive device for a cleaning robot provided in an embodiment of this application;

[0091] Figure 4 This is a schematic diagram of the electronic device structure shown in an embodiment of this application. Detailed Implementation

[0092] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0093] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the application. The singular forms “a,” “said,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms, and “multiple” generally includes at least two unless the context clearly indicates otherwise.

[0094] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0095] It should be understood that although the terms first, second, third, etc., may be used in the embodiments of this application, these descriptions should not be limited to these terms. These terms are only used to distinguish the descriptions. For example, first may also be referred to as second without departing from the scope of the embodiments of this application, and similarly, second may also be referred to as first.

[0096] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that an article or device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such an article or device. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the article or device that includes said element.

[0097] The optional embodiments of this application are described in detail below with reference to the accompanying drawings.

[0098] Figure 1A flowchart illustrating an intelligent interaction method for a cleaning robot provided in this application embodiment is shown below. Figure 1 As shown, the intelligent interaction method of this cleaning robot includes the following steps:

[0099] S101. In response to the video communication start command, establish a connection with the communication peer;

[0100] Establishing a connection with the communication peer refers to establishing a video communication session with the communication peer, that is, successfully initiating and establishing a two-way audio and video data stream transmission link.

[0101] The communication peer can be a mobile phone, tablet, computer, or other electronic device that can remotely control the cleaning robot and has video receiving capabilities, and includes a cleaning robot application (APP).

[0102] Understandably, the initiator of the video communication start command can be a local user in the same space as the cleaning robot, or a peer user in the communication field. This peer user can be a remote user (or a local user in the same space as the cleaning robot).

[0103] S102, Based on pre-stored features in the environment , Identify the local communication target in video communication;

[0104] Understandably, the pre-stored features are the characteristics of a pre-defined whitelist of users authorized to conduct video communication. These features can be biometrics. The video communication initiation command can identify the identities of both parties in the video communication. Then, based on the pre-stored features (biometrics) of the local communication partner, the system searches the environment for a matching user. For example, if the local communication partner identified from the video communication command is a child at home, the system retrieves the child's pre-stored features and searches the environment for a user whose features match. Once found, the user can be accompanied for a short distance to further verify their identity and avoid misidentification.

[0105] S103. When performing a video communication task, actively accompany the local communication object at a preset accompanying distance, and ensure that the local communication object is in the video frame transmitted to the communication peer.

[0106] It is worth noting that the active companionship here refers to the cleaning robot moving within a suitable distance (such as 1-2 meters) directly in front of or diagonally in front of the local communication object, adjusting its own position and shooting angle in real time to ensure that the user is always in a reasonable area of ​​the video frame, so that the communication counterpart can see the user clearly and continuously.

[0107] In this context, the local communication object is located in the video frame transmitted to the communication peer. This can be either the upper body of the local communication object being in the video frame transmitted to the communication peer (i.e., only the upper body is transmitted), or the entire body of the local communication object being in the video frame transmitted to the communication peer.

[0108] In this embodiment, a connection is established with the communication peer in response to a video communication initiation command. After the connection is established, the local communication object is identified based on pre-stored features in the environment, thus formally initiating the video communication task between the local communication object and the communication peer. During the execution of the video communication task, the robot actively accompanies the local communication object at a preset accompaniment distance, ensuring that the local communication object is within the video frame transmitted to the communication peer. Therefore, the cleaning robot is no longer a fixed monitoring probe, but an intelligent videographer that actively accompanies the user's movement. It can automatically lock onto the video communication object and maintain an appropriate distance for active accompaniment, ensuring that the other party is always clearly visible in the frame. This solves the problem of video frame loss or interruption caused by the user's (local communication object's) movement, making video communication while in motion continuous and stable.

[0109] Figure 2 A flowchart illustrating an intelligent interaction method for a cleaning robot provided in this application embodiment is shown below. Figure 2 As shown, the intelligent interaction method of this cleaning robot includes the following steps:

[0110] S201. In response to the video communication start command, establish a connection with the communication peer;

[0111] In some embodiments, the video communication activation command includes at least one of the following:

[0112] Receive video communication commands input by the user through the operation interface of the cleaning robot itself; for example, the local user can directly operate the robot's physical buttons, touch screen, and other interfaces to manually initiate video communication.

[0113] The cleaning robot's local voice interaction module recognizes video communication commands issued by the user; for example, when the local user is within the robot's voice recognition range, they can trigger video communication by giving verbal commands (such as saying "start video communication", "call mom", "call user number 2", etc. to the cleaning robot).

[0114] Receive video communication establishment request instructions from terminal devices; for example, users can send video communication requests remotely or locally through terminal apps such as mobile phones and tablets that are bound to the cleaning robot (such as clicking "Call Robot" or initiating a video connection on the mobile app), and the local user can establish a communication connection between the cleaning robot and the terminal device.

[0115] After receiving a call from a preset contact, the robot generates and executes a video communication start command. For example, if a preset contact (such as a family member) initiates a video call through a terminal device, the robot will automatically start video communication after receiving the call. No additional operation is required from the local user. If the caller is on the robot's pre-stored whitelist, a video communication connection will be automatically established.

[0116] S202. Based on the video communication start command, determine the identity information of the communication object on this end;

[0117] Understandably, parsing the video communication initiation command extracts crucial information about who wants to video chat with whom, enabling subsequent identification and proactive companionship. For example, a mother in the office app selects the "Baby" button from the family list to initiate a video call to the home cleaning robot. The command explicitly designates "Baby" as the local communication target, allowing the cleaning robot to access its pre-stored "Baby" feature database. Similarly, if the command is a voice command containing a title, such as "Call Dad," then "Dad" is clearly the remote communication target, and the voice initiator is the local communication target. The local user who issued the voice command is identified as the local communication target for this video call, and the robot prepares to use the user's pre-stored or real-time features for identification and proactive companionship.

[0118] S203. Based on the identity information, retrieve the pre-stored features corresponding to the identity information from the pre-stored biometric database;

[0119] Understandably, the pre-stored features are the characteristics of a pre-defined whitelist of users authorized to conduct video communication. These features can be biometrics. The video communication initiation command can identify the identities of both parties in the video communication. Then, based on the pre-stored features (biometrics) of the local communication partner, the system searches the environment for a matching user. For example, if the local communication partner identified from the video communication command is a child at home, the system retrieves the child's pre-stored features and searches the environment for a user whose features match. Once found, the user can be accompanied for a short distance to further verify their identity and avoid misidentification.

[0120] S204. Detect humanoid targets in the environment and collect multimodal biometrics of each humanoid target;

[0121] In some embodiments, multimodal biometrics include at least two of image features, voiceprint features, and gait features; wherein image features include facial features or human skeletal keypoint features.

[0122] Understandably, the image acquisition module of a cleaning robot has a relatively low line of sight. Since the module cannot perform pitching movements, it's easy to capture the user's legs when the robot is close to them. Therefore, gait characteristics are a relatively easy and important feature to obtain for quickly locating the target. These image features can be facial features, or upper body, lower body, or full-body image features.

[0123] In some embodiments, detecting humanoid targets in the environment includes:

[0124] Control the cleaning robot to move within the environment and collect environmental data in real time;

[0125] Detect human-shaped targets in the environment based on human silhouette features contained in environmental data.

[0126] For example, environmental data can be collected using environmental sensors built into the cleaning robot, such as environmental images captured by a camera, distance information collected by a LiDAR / TOF sensor, and thermal imaging information collected by an infrared sensor. Humanoid targets in the environment can then be detected using target detection algorithms (such as YOLO and SSD), contour extraction algorithms, human morphology recognition, and skeleton detection technologies.

[0127] This embodiment combines image, voiceprint, and gait features to avoid incomplete acquisition or recognition errors from a single feature, resulting in more accurate identification of the communicating object. Addressing the robot's low line of sight and inability to pitch, even with only leg information collected, gait recognition is possible, overcoming hardware limitations. By acquiring facial, upper and lower body, or full-body data, adjustments can be made flexibly based on distance and occlusion, ensuring effective feature acquisition in various scenarios.

[0128] S205. The pre-stored features are fused and matched with the multimodal biometric features of each humanoid target, and the local communication object of the video communication is identified based on the matching result.

[0129] The cleaning robot, once activated, uses its camera to locate people in the room (detecting humanoid targets) and captures their facial images, audio clips, or leg movements (extracting biometrics). Then, it compares the characteristics of each person captured on-site with pre-stored characteristics, calculating the similarity. If a person's (e.g., a child's) characteristics highly match the "baby's" pre-stored characteristics, a match is successful. The cleaning robot then confirms that this child is its target for local communication. Once confirmed, the robot actively accompanies the child and stably transmits their image to a remote video communication target (e.g., the mother).

[0130] S206. When performing a video communication task, actively accompany the local communication object at a preset accompanying distance, and ensure that the local communication object is in the video frame transmitted to the communication peer.

[0131] Understandably, when performing video communication tasks, the cleaning robot will actively follow the communication target (such as a child) at a preset safe distance (a walking distance, such as 1-2 meters, which will not disturb the user's activities and can stably record). At the same time, by adjusting its own position and camera angle, it will always keep the communication target in a reasonable area of ​​the video frame (such as the center or core field of view), ensuring that the communication recipient (such as a mother on a business trip or a family member far away) can see the target clearly and continuously, and avoid loss of image or unbalanced composition due to user movement.

[0132] In some embodiments, actively accompanying the local communication object at a preset accompanying distance and placing the local communication object in the video frame transmitted to the communication peer includes:

[0133] Based on the video stream of video communication, adjust the position of the cleaning robot and / or the orientation of the cleaning robot's image acquisition module to maintain the frontal image of the communicating object at the preset position on the display screen of the communicating end.

[0134] By analyzing the real-time video stream of video communication, the cleaning robot can automatically adjust its own position or the orientation of the camera (such as rotation angle or pitch angle) to ensure that the face image of the local caller is always in the preset area (such as the center of the screen) on the remote user's screen.

[0135] For example, the cleaning robot includes an image acquisition module and a gimbal mechanism. The image acquisition module is located on the gimbal mechanism. The gimbal structure can adjust the acquisition angle (such as horizontal rotation angle, pitch angle), shooting height and shooting orientation of the image acquisition module, so as to achieve accurate framing and tracking of the communication object on this end, and ensure that its frontal image is stably placed in the preset position of the display screen on the communication end.

[0136] In some embodiments, adjusting the position of the cleaning robot and / or the orientation of the cleaning robot's image acquisition module includes:

[0137] Based on the video stream from video communication, the facial orientation of the communicating object at this end is obtained;

[0138] Based on the deviation between the face orientation and the facing direction, adjust the position of the cleaning robot and / or the orientation of the image acquisition module so that the frontal image of the object being communicated on this end is located at the preset display position on the other end of the communication.

[0139] The "direction facing" refers to the center shooting direction of the image acquisition module (such as a camera) of the cleaning robot, that is, the straight line direction pointed to by the optical axis of the lens when the image acquisition module is in the current shooting posture.

[0140] For example, facial images of the communicating object are extracted in real time from the video stream, and key facial feature points (such as the center of the pupil, the corner of the eye, the tip of the nose, the corner of the mouth, etc.) of the communicating object are extracted from the facial images to obtain the facial orientation in three dimensions.

[0141] In this embodiment, through precise recognition of facial orientation and gaze direction, the cleaning robot and image acquisition module are adjusted to better match the user's actual viewing habits. This avoids the user's face deviating from the frame due to slight head turns or tilts, ensuring the image seen by the other end of the communication is more realistic and the interaction is more natural. Targeted adjustments based on quantified deviation data prevent image jitter and shifts caused by blind movements or angle adjustments, ensuring the user's face remains consistently in the preset display position (e.g., center of the frame) on the other end. This reduces visual fatigue for the other end and improves the continuity and clarity of video communication.

[0142] In some embodiments, adjusting the position of the cleaning robot and / or the orientation of the cleaning robot's image acquisition module includes:

[0143] Identify the current position of the face of the local communication target in the video frame;

[0144] Calculate the offset between the current position and the preset position on the local screen;

[0145] Adjust the position of the cleaning robot and / or the orientation of the image acquisition module according to the offset, so that the frontal image of the object being communicated on this end is located at the preset display position on the other end.

[0146] Understandably, the preset reference position (such as the center area of ​​the screen) in the local image of the cleaning robot is a mirror reference of the preset display position (such as the center of the screen on the other end). That is, when the user's face is in the preset position in the local image, the image received by the other end after video stream transmission naturally conforms to the requirement that the face is in the preset area (the transmission process only performs image quality compression and does not change the relative position of the target in the image). Therefore, the local image meeting the standard is a prerequisite for the end image meeting the standard. Thus, based on the current position of the face of the communicating object in the local video frame, the position of the cleaning robot and / or the orientation of the image acquisition module are adjusted to ensure that the face image of the communicating object is located in the preset display position on the other end.

[0147] In some embodiments, adjusting the position of the cleaning robot and / or the orientation of the image acquisition module so that the frontal image of the object being communicated on this end is located at a preset display position on the other end of the communication, including:

[0148] When it is determined that the cleaning robot cannot move, the orientation of the image acquisition module is adjusted so that the frontal image of the object being communicated with is located at the preset display position on the other end of the communication; or,

[0149] When determining the movable position of the cleaning robot, the priority of the image acquisition module orientation and the cleaning robot position adjustment is determined based on the energy consumption of the cleaning robot position adjustment and the image acquisition module orientation adjustment.

[0150] Understandably, before adjusting the position of the cleaning robot and / or the orientation of the image acquisition module, the feasibility of the robot's movement is first determined based on environmental data (such as obstacle information collected by LiDAR and cameras, and spatial boundary data). For example, if there are insurmountable obstacles such as walls or tall cabinets in a suitable shooting position that allows the user to face the camera directly, or if the position is beyond the robot's movement range (such as room boundaries), then the robot is determined to be unable to move. If the target shooting position is unobstructed and within the movement range, then the robot is determined to be able to move.

[0151] If the position cannot be moved, the orientation of the image acquisition module is directly adjusted (e.g., by using a gimbal mechanism to achieve horizontal rotation and pitch angle fine-tuning) to maximize the position of the frontal image of the communicating object at the preset display position on the other end, thus avoiding collisions and jamming caused by the robot moving blindly.

[0152] If it is determined that the position can be moved, the energy consumption difference between adjusting only the orientation of the image acquisition module and adjusting the robot position (either alone or in conjunction with angle adjustment) is further compared (e.g., the energy consumption of gimbal fine-tuning is much lower than the energy consumption of moving the robot's drive wheels). The adjustment priority is determined, and the method with lower energy consumption is selected first (usually adjusting the orientation of the image acquisition module alone). If angle adjustment alone cannot meet the image requirements (e.g., the user's offset distance is too large, and the gimbal angle has reached its limit), then the principle of low energy consumption is prioritized, and small-amplitude position adjustments and precise angle fine-tuning are performed in conjunction to ensure that the image on the other end meets the standards while reducing the overall energy consumption of the robot.

[0153] In some embodiments, actively accompanying the local communication object at a preset accompanying distance and placing the local communication object in the video frame transmitted to the communication peer includes:

[0154] When the facial image of the local communication object fails to meet the preset communication quality conditions, an image frame containing a visually enhanced facial image is generated based on the pre-stored facial feature information of the local communication object, and then output to the other end of the communication.

[0155] Understandably, communication quality conditions include ensuring that the clarity of the facial image of the local communication subject transmitted to the peer device meets the requirement of clearly discernible facial expressions, so as to ensure that the peer communication subject can accurately recognize the facial expressions of the local communication subject (such as smiling, nodding, etc.) and thus guarantee the effectiveness of video interaction.

[0156] For example, failure to meet preset communication quality conditions includes at least one of the following situations:

[0157] Face loss; for example, the front of the local communication object can be captured, but the face is completely obscured; or, the local communication object is facing away from the image acquisition module, so a facial image cannot be captured.

[0158] The area of ​​the face that is obscured exceeds a preset threshold;

[0159] The face is not at a frontal angle; for example, due to the location of the cleaning robot and the orientation of the image acquisition module, only the side of the face can be captured, not the front of the face.

[0160] The facial image clarity is below the preset threshold.

[0161] In some embodiments, generating an image frame that includes a visually enhanced facial image includes:

[0162] Generate a fitted facial image that matches the current head pose and body orientation of the communicating object on this end;

[0163] The fitted facial image is fused with the background of the current image frame to generate and output an image frame containing visually enhanced facial images.

[0164] Among them, matching the current head posture and body orientation of the communicating object means that the generated fitted frontal facial image is highly consistent with the head posture and body orientation of the communicating object in terms of spatial logic and visual presentation. In other words, the fitted image is not a standard frontal face generated out of thin air, but is based on the user's real-time head posture (tilt, twist) and body orientation (chest facing the direction) to ensure that the relationship between the head tilt angle and body orientation under the perspective of the human face facing forward conforms to the movement law of the real human body, thus presenting a natural and realistic frontal perspective.

[0165] In some embodiments, before fusing the fitted facial image with the background of the current image frame, the method further includes:

[0166] Evaluate the visual consistency coefficient between the fitted facial image and the body region of the local communication object in the current frame;

[0167] Based on visual coordination, facial fusion or body replacement is performed.

[0168] Understandably, facial image fitting and fusion are only performed when the human face is facing forward and the orientation of the body regions conforms to the laws of the real human body; otherwise, the current frame image can be directly replaced.

[0169] In some embodiments, facial fusion or body replacement is performed based on visual coordination, including:

[0170] When the visual coordination coefficient is higher than the preset threshold, the fitted facial image will be fused with the original body region in the current frame.

[0171] When the visual coordination coefficient is lower than or equal to a preset threshold, a pre-stored half-body or full-body reference image containing a frontal face is used to replace the body region of the local communication object in the current frame and blend it with the background of the current image frame.

[0172] Among them, a value greater than the preset threshold indicates that the fitted face and body regions conform to the laws of real human movement. However, a value less than or equal to the preset threshold indicates that there is a visual discrepancy between the fitted face and body regions (such as a serious disconnect between the fitted face orientation and the torso pointing). In this case, the face fitting is abandoned, the current image frame is discarded, and the current image frame is replaced.

[0173] This embodiment avoids the problem of side / half-side faces appearing in the original video stream due to the communicating party turning their head left or right, looking down, or tilting their head, making it difficult for the communicating end to clearly recognize facial features and expressions (such as smiling, frowning, nodding, etc.). By generating a fitted frontal face that matches the current posture, it ensures that the communicating end can always obtain a clear and complete facial image, accurately capture emotional feedback and non-verbal communication information, and improve the interactive efficiency of video communication.

[0174] In some embodiments, the method further includes: during video communication, responding to an active companion behavior mode switching instruction, and switching the current active companion behavior from a first active companion mode to a second active companion mode according to the active companion behavior mode switching instruction;

[0175] Among them, the first active companion mode and the second active companion mode differ in at least one of the active companion parameters, namely companion distance, active companion speed, camera focal length, and camera pitch angle.

[0176] In this embodiment, by switching the companion mode (adjusting parameters such as companion distance, speed, camera focal length, and tilt angle), different scenarios in video communication (such as close-up communication, long-distance panoramic interaction, and follow-up shooting while moving) can be matched, avoiding the limitation of a single mode being unable to adapt to multiple scenarios. Of course, this embodiment also supports users to switch modes actively via commands to meet different usage habits (such as some users preferring close-up views, while others prefer panoramic views), without having to manually adjust individual parameters, thus lowering the operational threshold and enhancing user control.

[0177] In some embodiments, the active companion parameters of the first active companion mode cause the face of the local communication object to be located in the center of the display screen of the communication peer.

[0178] The active companion mode's active companion parameters ensure that the entire body of the local communication object is centered on the display screen of the communication peer.

[0179] Understandably, the first active companion mode dialogue mode highlights facial details by centering the face, while the second active companion mode presents the environment and actions by centering the body, effectively solving the pain point that a single mode cannot take into account both close-ups and panoramic views.

[0180] In some scenarios, the first application of proactive companionship is as follows: when a mother is away on a business trip, she can have a bedtime video call with her child. The cleaning robot automatically shortens the companionship distance, increases the camera focal length, places the child's face in the center of the frame and highlights the details. The mother can clearly see the child's expression (such as whether he is happy or upset), and the child can also clearly see the mother's facial expression, achieving intimate communication like face-to-face communication and avoiding blurring of facial details due to too much irrelevant background in the picture.

[0181] The second application scenario for active companionship: When a child shows their mother newly learned handicrafts, dance moves, or painting achievements at home, the mother can switch to observation mode by actively commanding the robot. The robot will automatically increase the companionship distance and adjust the camera focus and tilt angle to present the child's whole body and surrounding environment (such as the craft table or dance area) in the center of the screen. This solves the problem that the single dialogue mode cannot show the whole body movements and scene details, making remote interaction more contextual and practical.

[0182] In some embodiments, the method further includes:

[0183] During video communication, the cleaning robot responds to the auxiliary task instructions and executes the auxiliary tasks corresponding to the instructions.

[0184] After the auxiliary task is completed, control the cleaning robot to adjust its own position and / or camera orientation so that the frontal image of the object being communicated with is located at the preset display position of the other end of the communication;

[0185] During the execution of ancillary tasks, a video communication connection with the communication peer is maintained; the ancillary tasks include at least one of the following:

[0186] Perform local cleaning at the current location of the local communication object;

[0187] Move to the preset location and collect environmental data at the preset location;

[0188] The cleaning robot's output module outputs preset audio or image content that is independent of video communication.

[0189] Among them, the auxiliary task is not an independent task, but a task that is executed synchronously when performing the video communication task.

[0190] In some scenarios, a mother might learn through a video call that her child has spilled milk on the living room carpet, and she can remotely instruct the cleaning robot to perform spot cleaning. While traveling, the mother can also use a video call to direct the cleaning robot to check the temperature and humidity in her child's room to ensure it's comfortable (e.g., whether it's stuffy in summer or dry in winter). Before ending the video call, the mother can instruct the cleaning robot to play preset audio (e.g., soothing music).

[0191] In some embodiments, the method further includes: when waiting for or executing a first independent task, and receiving a trigger instruction for a second independent task, determining a scheduling scheme for the first independent task and the second independent task based on a preset task arbitration decision model, according to the identity information of the user associated with the first independent task and the identity information of the initiator of the second independent task, as well as the preset execution priorities of the first independent task and the second independent task.

[0192] The scheduling scheme includes:

[0193] Interrupt the execution of the first independent task and switch to the second independent task;

[0194] Maintain execution of the first independent task, but do not execute the second independent task; or,

[0195] Send a confirmation request to the user associated with the first independent task, and switch to the second independent task or continue executing the first independent task based on the feedback.

[0196] When the cleaning robot is handling a task (the first independent task) and receives a new task instruction (the second independent task), it will decide how to arrange the two tasks according to the preset task arbitration rules.

[0197] In this embodiment, by using both identity information and task priority to determine the priority of urgent and important tasks (such as video communication initiated by the mother), the system ensures that secondary tasks (such as cleaning tasks) are executed first, thereby avoiding the consumption of resources by secondary tasks and solving the technical defect that important and urgent tasks cannot be met in a timely manner due to multi-task conflicts.

[0198] In some scenarios, the cleaning robot is performing the first independent task (initiated by the mother, collecting temperature and humidity data in the child's room, associated with the mother) when it receives a command from the child to play an animated audio clip (the second independent task, initiated by the child). Because the mother's identity level is higher than the child's, the arbitration model determines that the data collection should continue, and the task of playing the animation should not be executed for the time being.

[0199] In some embodiments, the decision rules of the task arbitration decision model include:

[0200] Based on the preset permission and identity levels of the user associated with the first independent task and the initiator of the second independent task, a scheduling scheme is determined.

[0201] When the permission levels of the identity information of the user associated with the first independent task and the identity information of the initiator of the second independent task are the same, a scheduling scheme is determined based on the preset execution priorities of the first and second independent tasks.

[0202] The process first determines task scheduling based on the user's preset permission level (e.g., the mother's permission level is higher than the child's). If the permission levels are different, tasks initiated by the user with higher permissions take priority. If both users have the same permission level (e.g., both are family members with ordinary permissions), the scheduling scheme is then determined based on the preset priority of the task itself (e.g., video communication task > cleaning task).

[0203] In this embodiment, identity permissions and task priority are used for dual determination, which makes the scheduling logic clearer. Priority is given to the needs of users with high permissions, while also taking into account the urgency of tasks with the same permissions. This avoids permission conflicts and disordered task execution, and improves the rationality and efficiency of scheduling.

[0204] In some embodiments, the scheduling scheme determined based on the preset execution priorities of the first independent task and the second independent task includes at least one of the following:

[0205] When the first independent task is a video communication task, output the scheduling scheme to maintain the execution of the first independent task;

[0206] When the preset execution priorities of the first independent task and the second independent task are the same, the output will send a confirmation query to the user associated with the first independent task, and switch to the second independent task or maintain the scheduling scheme of executing the first independent task based on the feedback result.

[0207] When the preset execution priorities of the first independent task and the second independent task are different, the scheduling scheme of the task with the higher execution priority is output.

[0208] In some embodiments, the first independent task and the second independent task are both any one of a cleaning task, a video communication task, and an interactive task; wherein, the interactive task includes: recording the behavior of an interactive object, and outputting at least one of motion feedback, sound feedback, light effect feedback, and screen feedback based on the behavior.

[0209] Understandably, the basic tasks of cleaning robots include spot cleaning and whole-house cleaning. Video communication tasks involve real-time video calls and motion companionship between the robot and a peer (such as a mother and child or family members). Interactive tasks involve the cleaning robot recording the behavior of the interacting object through sensors (cameras, microphones, etc.) and then outputting corresponding feedback.

[0210] In some embodiments, the objects of action interaction include humans and animals.

[0211] In one scenario, when a child waves, the cleaning robot outputs a "Hello~" sound as feedback.

[0212] In another scenario, the cleaning robot detects the cat approaching and activates its laser pointer function (preset cat-teasing mode), projecting laser dots along the cat's movement trajectory to guide the cat to interact.

[0213] This application also provides apparatus embodiments that follow the above embodiments, configured to implement the method steps of the above embodiments. The interpretation of the same names is the same as that of the above embodiments, and they have the same technical effects as those of the above embodiments, so they will not be described again here.

[0214] like Figure 3 As shown, this application provides an intelligent interactive device for a cleaning robot, the device comprising:

[0215] The communication connection unit 301 is configured to establish a connection with the communication peer in response to a video communication start command.

[0216] The target search unit 302 is configured to identify the local communication object of the video communication based on pre-stored features in the environment;

[0217] The companion control unit 303 is configured to actively accompany the local communication object at a preset companion distance when performing a video communication task, and to place the local communication object in the video frame transmitted to the communication peer.

[0218] In some embodiments, the accompanying control unit 303 is also configured to adjust the position of the cleaning robot and / or the orientation of the cleaning robot's image acquisition module based on the video stream of video communication, so as to maintain the frontal image of the local communication object at a preset position on the display screen of the communication peer.

[0219] In some embodiments, the accompanying control unit 303 is further configured to obtain the facial orientation of the local communication object based on the video stream of the video communication; and adjust the position of the cleaning robot and / or the orientation of the image acquisition module according to the deviation between the facial orientation and the facing direction, so that the frontal image of the local communication object is located at a preset display position of the communication counterpart.

[0220] In some embodiments, the companion control unit 303 is further configured to identify the current position of the face of the local communication object in a video frame; calculate the offset between the current position and a preset position on the local screen; and adjust the position of the cleaning robot and / or the orientation of the image acquisition module according to the offset, so that the frontal image of the local communication object is located at a preset display position on the other end of the communication.

[0221] In some embodiments, the accompanying control unit 303 is further configured to adjust the orientation of the image acquisition module when it is determined that the cleaning robot cannot move, so that the frontal image of the local communication object is located at a preset display position of the communication peer; or, when it is determined that the cleaning robot can move, determine the priority of the image acquisition module orientation and the cleaning robot position adjustment based on the energy consumption of the cleaning robot position adjustment and the image acquisition module orientation adjustment.

[0222] In some embodiments, the companion control unit 303 is further configured to generate an image frame containing a visually enhanced facial image based on pre-stored facial feature information of the local communication object when the facial image of the local communication object fails to meet preset communication quality conditions, and output it to the communication peer.

[0223] In some embodiments, the inability to meet the preset communication quality conditions includes at least one of the following situations:

[0224] Facial loss;

[0225] The area of ​​the face that is obscured exceeds a preset threshold;

[0226] The face is not shown from a frontal angle;

[0227] The facial image clarity is below the preset threshold.

[0228] In some embodiments, the companion control unit 303 is further configured to generate a fitted facial image that matches the current head posture and body orientation of the local communication object;

[0229] The fitted facial image is fused with the background of the current image frame to generate and output an image frame containing visually enhanced facial images.

[0230] In some embodiments, the companion control unit 303 is further configured to evaluate the visual coordination coefficient between the fitted facial image and the body region of the local communication object in the current frame;

[0231] Based on visual coordination, facial fusion or body replacement is performed.

[0232] In some embodiments, the companion control unit 303 is further configured to fuse the fitted facial image with the original body region in the current frame when the visual coordination coefficient is higher than a preset threshold; and to replace the body region of the local communication object in the current frame with a pre-stored half-body or full-body reference image containing a frontal face when the visual coordination coefficient is lower than or equal to the preset threshold, and to fuse it with the background of the current image frame.

[0233] In some embodiments, the companion control unit 303 is further configured to respond to an active companion behavior mode switching instruction during video communication, and switch the current active companion behavior from a first active companion mode to a second active companion mode according to the active companion behavior mode switching instruction; wherein the first active companion mode and the second active companion mode differ in at least one active companion parameter among companion distance, active companion speed, camera focal length, and camera pitch angle.

[0234] In some embodiments, the active companion parameters of the first active companion mode position the face of the local communication object at the center of the display screen of the communication peer; the active companion parameters of the second active companion mode position the entire body of the local communication object at the center of the display screen of the communication peer.

[0235] In some embodiments, the video communication activation command includes at least one of the following:

[0236] Receives video communication commands input by the user through the operation interface of the cleaning robot itself;

[0237] The cleaning robot's local voice interaction module recognizes video communication commands issued by the user.

[0238] Receive a video communication establishment request command from the terminal device;

[0239] Upon receiving a call from a preset contact, a video communication initiation command is generated and executed.

[0240] In some embodiments, the target search unit 302 is further configured to determine the identity information of the local communication object based on the video communication start instruction; retrieve the pre-stored features corresponding to the identity information from a pre-stored biometric database based on the identity information; and identify the local communication object of the video communication based on the matching result of the pre-stored features and the biometric features of humanoid targets detected in the environment.

[0241] Based on the matching results, the local communication object of the video communication is identified.

[0242] In some embodiments, the target search unit 302 is further configured to detect humanoid targets in the environment and collect multimodal biometrics of each humanoid target; fuse and match pre-stored features with the multimodal biometrics of each humanoid target, and identify the local communication object of the video communication based on the matching result.

[0243] In some embodiments, the target search unit 302 is also configured to control the cleaning robot to move within the environment and collect environmental data in real time; and to detect humanoid targets in the environment based on human contour features contained in the environmental data.

[0244] In some embodiments, the target search unit 302 is further configured to use multimodal biometrics including at least two of image features, voiceprint features, and gait features; wherein the image features include facial features or human skeletal keypoint features.

[0245] In some embodiments, the companion control unit 303 is further configured to, during video communication, the cleaning robot responds to the auxiliary task instruction and executes the auxiliary task corresponding to the auxiliary task instruction; after the auxiliary task is executed, the cleaning robot is controlled to adjust its own position and / or camera orientation so that the frontal image of the communication object is located at the preset display position of the communication counterpart.

[0246] During the execution of ancillary tasks, a video communication connection with the communication peer is maintained; the ancillary tasks include at least one of the following:

[0247] Perform local cleaning at the current location of the local communication object;

[0248] Move to the preset location and collect environmental data at the preset location;

[0249] The cleaning robot's output module outputs preset audio or image content that is independent of video communication.

[0250] In some embodiments, the device further includes a task scheduling module, configured to, when waiting for or executing a first independent task and receiving a trigger instruction for a second independent task, determine a scheduling scheme for the first independent task and the second independent task based on a preset task arbitration decision model, according to the identity information of the user associated with the first independent task and the identity information of the initiator of the second independent task, as well as the preset execution priorities of the first independent task and the second independent task.

[0251] The scheduling scheme includes:

[0252] Interrupt the execution of the first independent task and switch to the second independent task;

[0253] Maintain execution of the first independent task, but do not execute the second independent task; or,

[0254] Send a confirmation request to the user associated with the first independent task, and switch to the second independent task or continue executing the first independent task based on the feedback.

[0255] In some embodiments, the decision rules of the task arbitration decision model include:

[0256] Based on the preset permission and identity levels of the user associated with the first independent task and the initiator of the second independent task, a scheduling scheme is determined.

[0257] When the permission levels of the identity information of the user associated with the first independent task and the identity information of the initiator of the second independent task are the same, a scheduling scheme is determined based on the preset execution priorities of the first and second independent tasks.

[0258] In some embodiments, the scheduling scheme determined based on the preset execution priorities of the first independent task and the second independent task includes at least one of the following:

[0259] When the first independent task is a video communication task, output the scheduling scheme to maintain the execution of the first independent task;

[0260] When the preset execution priorities of the first independent task and the second independent task are the same, the output will send a confirmation query to the user associated with the first independent task, and switch to the second independent task or maintain the scheduling scheme of executing the first independent task based on the feedback result.

[0261] When the preset execution priorities of the first independent task and the second independent task are different, the scheduling scheme of the task with the higher execution priority is output.

[0262] In some embodiments, the first independent task and the second independent task are both any one of a cleaning task, a video communication task, and an interactive task; wherein, the interactive task includes: recording the behavior of an interactive object, and outputting at least one of motion feedback, sound feedback, light effect feedback, and screen feedback based on the behavior.

[0263] In some embodiments, the objects of action interaction include humans and animals.

[0264] like Figure 4 As shown, this embodiment provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by a processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method steps of the above embodiment.

[0265] This application provides a non-volatile computer storage medium storing computer-executable instructions that can execute the method steps of the above embodiments.

[0266] The following is for reference. Figure 4 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of this application. The terminal devices in the embodiments of this application may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0267] like Figure 4 As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. The RAM 403 also stores various programs and data required for the operation of the electronic device. The processing unit 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0268] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 408 including, for example, magnetic tapes, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0269] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code configured to perform the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 409, or installed from storage device 408, or installed from ROM 402. When the computer program is executed by processing device 401, it performs the functions defined in the methods of embodiments of this application.

[0270] It should be noted that the computer-readable medium described above in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program configured for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0271] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0272] Computer program code configured to perform the operations of this application can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0273] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions configured to implement a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0274] The units described in the embodiments of this application can be implemented in software or hardware. The names of the units are not, in some cases, limiting the scope of the unit itself.

Claims

1. An intelligent interaction method for a cleaning robot, characterized in that, The method includes: In response to a video communication initiation command, establish a connection with the communication peer; Based on pre-stored features in the environment, identify the local communication object in video communication; When performing a video communication task, the local communication object is actively accompanied at a preset accompanying distance, and the local communication object is placed in the video frame transmitted to the communication peer. The step of actively accompanying the local communication object at a preset accompanying distance and placing the local communication object in the video frame transmitted to the communication peer includes: When it is determined that the cleaning robot cannot move, the orientation of the image acquisition module is adjusted based on the video stream from the video communication, so that the frontal image of the local communication object is located at a preset display position on the other end of the communication; or... When determining the movable position of the cleaning robot, the priority of the orientation of the image acquisition module and the position adjustment of the cleaning robot is determined based on the video stream of the video communication, the separate energy consumption of the cleaning robot's position adjustment, and the separate energy consumption of the orientation adjustment of the image acquisition module. According to the adjustment method corresponding to the priority, the frontal image of the local communication object is placed at the preset display position of the communication peer.

2. The method according to claim 1, characterized in that, Adjusting the position of the cleaning robot and / or the orientation of the image acquisition module of the cleaning robot includes: Based on the video stream from video communication, the facial orientation of the local communication object is obtained; Based on the deviation between the facial orientation and the facing direction, adjust the position of the cleaning robot and / or the orientation of the image acquisition module so that the frontal image of the local communication object is located at the preset display position of the communication peer.

3. The method according to claim 2, characterized in that, Adjusting the position of the cleaning robot and / or the orientation of the image acquisition module of the cleaning robot includes: Identify the current position of the face of the local communication object in the video frame; Calculate the offset between the current position and the preset position on the local screen; Based on the offset, adjust the position of the cleaning robot and / or the orientation of the image acquisition module so that the frontal image of the local communication object is located at the preset display position of the communication peer.

4. The method according to claim 1, characterized in that, The step of actively accompanying the local communication object at a preset accompanying distance and placing the local communication object in the video frame transmitted to the communication peer includes: When the facial image of the local communication object fails to meet the preset communication quality conditions, an image frame containing a visually enhanced facial image is generated based on the pre-stored facial feature information of the local communication object, and output to the communication peer.

5. The method according to claim 4, characterized in that, The inability to meet the preset communication quality conditions includes at least one of the following situations: Facial loss; The area of ​​the face that is obscured exceeds a preset threshold; The face is not shown from a frontal angle; The facial image clarity is below the preset threshold.

6. The method according to claim 4, characterized in that, The process of generating image frames containing visually enhanced facial images includes: Generate a fitted facial image that matches the current head posture and body orientation of the local communication object; The fitted facial image is fused with the background of the current image frame to generate and output the image frame containing the visually enhanced facial image.

7. The method according to claim 6, characterized in that, Before fusing the fitted facial image with the background of the current image frame, the method further includes: Evaluate the visual consistency coefficient between the fitted facial image and the body region of the local communication object in the current frame; Based on the visual coordination, facial fusion or body replacement is performed.

8. The method according to claim 7, characterized in that, The process of performing facial fusion or body replacement based on the visual coordination includes: When the visual coordination coefficient is higher than a preset threshold, the fitted facial image is fused with the original body region in the current frame; When the visual coordination coefficient is lower than or equal to a preset threshold, a pre-stored half-body or full-body reference image containing a frontal face is used to replace the body region of the local communication object in the current frame and is blended with the background of the current image frame.

9. The method according to claim 1, characterized in that, The method further includes: during video communication, responding to an active companion behavior mode switching instruction, and switching the current active companion behavior from a first active companion mode to a second active companion mode according to the active companion behavior mode switching instruction; Among them, the first active companion mode and the second active companion mode differ in at least one of the active companion parameters, namely, companion distance, active companion speed, camera focal length, and camera pitch angle.

10. The method according to claim 9, characterized in that, The active companion mode of the first active companion mode positions the face of the local communication object at the center of the display screen of the communication peer; the active companion mode of the second active companion mode positions the entire body of the local communication object at the center of the display screen of the communication peer.

11. The method according to claim 1, characterized in that, The video communication activation command includes at least one of the following: Receive video communication commands input by the user through the operation interface of the cleaning robot body; The cleaning robot's local voice interaction module recognizes video communication commands issued by the user. Receive a video communication establishment request command from the terminal device; Upon receiving a call from a preset contact, a video communication initiation command is generated and executed.

12. The method according to claim 1, characterized in that, The process of identifying the local communication object in video communication based on pre-stored features in the environment includes: Based on the video communication start command, the identity information of the communication object on this end is determined; Based on the identity information, retrieve the pre-stored features corresponding to the identity information from the pre-stored biometric database; Based on the matching results between the pre-stored features and the biometric features of humanoid targets detected in the environment, the local communication object of the video communication is identified; Based on the matching results, the local communication object of the video communication is identified.

13. The method according to claim 1, characterized in that, The process of identifying the local communication target in video communication based on the matching results between the pre-stored features and the biometric features of humanoid targets detected in the environment includes: Detect humanoid targets in the environment and collect multimodal biometrics of each humanoid target; The pre-stored features are fused and matched with the multimodal biometric features of each humanoid target, and the local communication object of the video communication is identified based on the matching result.

14. The method according to claim 13, characterized in that, The humanoid targets in the detection environment include: Control the cleaning robot to move within the environment and collect environmental data in real time; Based on the human silhouette features contained in the environmental data, human-shaped targets in the environment are detected.

15. The method according to claim 13, characterized in that, The multimodal biometrics include at least two of image features, voiceprint features, and gait features; wherein the image features include facial features or human skeletal key point features.

16. The method according to claim 1, characterized in that, The method further includes: During video communication, the cleaning robot responds to ancillary task instructions and executes the ancillary tasks corresponding to those instructions. After the auxiliary task is completed, the cleaning robot is controlled to adjust its own position and / or camera orientation so that the frontal image of the local communication object is located at the preset display position of the communication counterpart. During the execution of the ancillary tasks, a video communication connection with the communication peer is maintained; the ancillary tasks include at least one of the following: Perform local cleaning at the current location of the local communication object; Move to a preset location and collect environmental data at that preset location; The output module of the cleaning robot is used to output preset audio or image content that is independent of the video communication.

17. The method according to claim 1, characterized in that, The method further includes: when waiting for or executing a first independent task, and receiving a trigger instruction for a second independent task, determining a scheduling scheme for the first independent task and the second independent task based on a preset task arbitration decision model, according to the identity information of the user associated with the first independent task and the identity information of the initiator of the second independent task, as well as the preset execution priorities of the first independent task and the second independent task. The scheduling scheme includes: Interrupt the execution of the first independent task and switch to the second independent task; Maintain execution of the first independent task, but do not execute the second independent task; or, Send a confirmation request to the user associated with the first independent task, and switch to the second independent task or continue executing the first independent task based on the feedback.

18. The method according to claim 17, characterized in that, The decision rules of the task arbitration decision model include: Based on the preset permission and identity levels of the user associated with the first independent task and the initiator of the second independent task, a scheduling scheme is determined. When the permission levels of the identity information of the user associated with the first independent task and the identity information of the initiator of the second independent task are the same, a scheduling scheme is determined based on the preset execution priorities of the first independent task and the second independent task.

19. The method according to claim 18, characterized in that, The process of determining a scheduling scheme based on the preset execution priorities of the first independent task and the second independent task includes at least one of the following: When the first independent task is a video communication task, output a scheduling scheme to maintain the execution of the first independent task; When the preset execution priorities of the first independent task and the second independent task are the same, a confirmation query is sent to the user associated with the first independent task, and the scheduling scheme of switching to the second independent task or maintaining the execution of the first independent task is switched according to the feedback result. When the preset execution priorities of the first independent task and the second independent task are different, the scheduling scheme of the task with the higher execution priority is output.

20. The method according to claim 19, characterized in that, Both the first independent task and the second independent task are any one of a cleaning task, a video communication task, and an interactive task; wherein, the interactive task includes: recording the behavior of an interactive object, and outputting at least one of motion feedback, sound feedback, light effect feedback, and screen feedback based on the behavior.

21. The method according to claim 20, characterized in that, The objects for interaction include humans and animals.

22. An intelligent interactive device for a cleaning robot, characterized in that, The device includes: The communication connection unit is configured to establish a connection with the communication peer in response to a video communication start command. The target search unit is configured to identify the local communication object of the video communication based on pre-stored features in the environment; The accompanying control unit is configured to actively accompany the local communication object at a preset accompanying distance when performing a video communication task, and to place the local communication object in the video frame transmitted to the communication peer. The step of actively accompanying the local communication object at a preset accompanying distance and placing the local communication object in the video frame transmitted to the communication peer includes: When it is determined that the cleaning robot cannot move, the orientation of the image acquisition module is adjusted based on the video stream from the video communication, so that the frontal image of the local communication object is located at a preset display position on the other end of the communication; or... When determining the movable position of the cleaning robot, the priority of the orientation of the image acquisition module and the position adjustment of the cleaning robot is determined based on the video stream of the video communication, the separate energy consumption of the cleaning robot's position adjustment, and the separate energy consumption of the orientation adjustment of the image acquisition module. According to the adjustment method corresponding to the priority, the frontal image of the local communication object is placed at the preset display position of the communication peer.

23. A cleaning robot, characterized in that, Including the intelligent interactive device as described in claim 22, the intelligent interactive device is configured to: In response to a video communication initiation command, establish a connection with the communication peer; Based on pre-stored features in the environment, identify the local communication object in video communication; When performing a video communication task, the local communication object is actively accompanied at a preset accompanying distance, and the local communication object is placed in the video frame transmitted to the communication peer. The step of actively accompanying the local communication object at a preset accompanying distance and placing the local communication object in the video frame transmitted to the communication peer includes: When it is determined that the cleaning robot cannot move, the orientation of the image acquisition module is adjusted based on the video stream from the video communication, so that the frontal image of the local communication object is located at a preset display position on the other end of the communication; or... When determining the movable position of the cleaning robot, the priority of the orientation of the image acquisition module and the position adjustment of the cleaning robot is determined based on the video stream of the video communication, the separate energy consumption of the cleaning robot's position adjustment, and the separate energy consumption of the orientation adjustment of the image acquisition module. According to the adjustment method corresponding to the priority, the frontal image of the local communication object is placed at the preset display position of the communication peer.

24. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 21.

25. An electronic device, characterized in that, include: One or more processors; A storage device configured to store one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 21.

Citation Information

Patent Citations

  • Robot and human face following method of robot

    CN109955248A

  • Person identity recognition method and system, and device and storage medium

    WO2024108606A1