Robot awakening method, device, computer equipment, readable storage medium and program product

By obtaining environmental pictures for face recognition and rotation control, the problem of waste of resources and awakening difficulties in traditional robot awakening methods is solved, and the robot is precisely awakened.

CN119992631BActive Publication Date: 2025-08-15SHANGHAI FOURIER INTELLIGENCE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510473556.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-08-15
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

Traditional robot awakening methods require specific wake-up words and different types of robot awakening words are different, resulting in waste of resources and difficulty in waking up specific robots accurately.

Method used

By obtaining the environmental pictures collected by the robot, facial recognition is performed, the face position is determined, and the robot rotation is controlled based on the position, so that the face is within the narrow beam sound range to wake up the robot, and only the sound in the narrow beam range is captured.

Benefits of technology

Accurate awakening of the robot is achieved, avoiding waste of resources and improving the accuracy and efficiency of awakening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992631B_ABST
    Figure CN119992631B_ABST
Patent Text Reader

Abstract

This application relates to a robot wake-up method, apparatus, computer device, computer-readable storage medium, and computer program product. The method comprises: obtaining environmental images captured by a robot; performing facial recognition on the environmental images; determining the location of a human face if a face is recognized in at least one of the environmental images; and controlling the robot's rotation based on the face's location so that the face is within the robot's narrow beam reception range, thereby waking the robot. This method can accurately wake up the robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of robot control technology, and in particular to a robot awakening method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Art

[0002] With the rapid development of modern technology, the application scope of intelligent robots is becoming more and more extensive. Intelligent robots can be seen in homes, shopping malls, banks and other public places.

[0003] However, in traditional technology, waking up a robot requires a specific wake-up word. Different types of robots have different wake-up words, while robots of the same type have the same wake-up word. In this way, when the same robot is used, many robots can be woken up by the same wake-up word, resulting in a waste of robot resources. In addition, if different types of robots are used, each robot requires a different wake-up word, and remembering these wake-up words is also challenging. Therefore, there is an urgent need for a way to accurately wake up the robot. Summary of the Invention

[0004] Based on this, it is necessary to provide a robot wake-up method, device, computer equipment, computer-readable storage medium and computer program product that can accurately wake up the robot in order to solve the above technical problems.

[0005] In a first aspect, the present application provides a robot awakening method, the method comprising:

[0006] Get the environment pictures collected by the robot;

[0007] Performing face recognition on the environment picture;

[0008] When a human face is recognized in at least one of the environment images, determining a position of the human face;

[0009] The robot is controlled to rotate based on the position of the human face so that the human face is located within the narrow beam receiving range of the robot to wake up the robot.

[0010] In one embodiment, determining the position of the face includes:

[0011] When at least one face is identified, the position of the face is determined in parallel;

[0012] Before controlling the rotation of the robot based on the position of the human face, the method further includes:

[0013] The determined positions of the human faces are sorted based on the relative distance between the narrow beam receiving range of the robot and the positions of the human faces.

[0014] In one embodiment, after determining the position of the face, the method further includes:

[0015] Segmenting the face image from the environment image based on the position of the face;

[0016] Performing mouth opening and closing detection on the face image;

[0017] If the face image contains an open or closed mouth, continuing to execute the step of controlling the rotation of the robot based on the position of the face;

[0018] When the face image does not have an open or closed mouth, the position of the face is ignored.

[0019] In one embodiment, after controlling the rotation of the robot based on the position of the human face, the method further comprises:

[0020] Real-time acquisition of the wake-up person's location where the face is within the robot's narrow beam reception range;

[0021] When the position of the awakening person changes, obtaining a first relative posture change of the awakening person;

[0022] determining a second relative posture change of the robot based on the first relative posture change;

[0023] The robot is controlled to track and wake up the person based on the change of the second relative posture of the robot.

[0024] In one embodiment, the robot includes at least two cameras, and the field of view of the at least two cameras constitutes the field of view of the robot; performing face recognition on the environment image includes:

[0025] Performing parallel face recognition on the environmental images captured by the cameras, and determining environmental images containing human faces;

[0026] The controlling the rotation of the robot based on the position of the human face so that the human face is located within the narrow beam receiving range of the robot to wake up the robot includes:

[0027] The robot is controlled to rotate in sequence based on the position of the human face so that the human face is located within the narrow beam sound receiving range of the robot to wake up the robot, and the awakened person is determined based on the acquired audio signal.

[0028] In one embodiment, after controlling the rotation of the robot based on the position of the human face, the method further comprises:

[0029] Get the audio signal received by the robot;

[0030] Converting the audio signal into text information;

[0031] Calling the large model to process the text information to obtain response information corresponding to the audio signal;

[0032] The response information is output.

[0033] In a second aspect, the present application further provides a robot awakening device, the device comprising:

[0034] Environmental image acquisition module, used to obtain environmental images collected by the robot;

[0035] A face recognition module, used to perform face recognition on the environment image;

[0036] a face position determination module, configured to determine the position of a face when a face is recognized in at least one of the environment images;

[0037] The control module is used to control the rotation of the robot based on the position of the human face so that the human face is located within the narrow beam receiving range of the robot to wake up the robot.

[0038] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method in any one of the above embodiments when executing the computer program.

[0039] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the method in any one of the above-mentioned embodiments when the computer program is executed by a processor.

[0040] In a fifth aspect, the present application also provides a computer program product, comprising a computer program, which implements the steps of the method in any one of the above embodiments when executed by a processor.

[0041] The above-mentioned robot wake-up method, device, computer equipment, computer-readable storage medium and computer program product can only capture the voice of a person speaking within the narrow-beam sound receiving range, and the voice of a person speaking outside the narrow-beam sound receiving range will not be captured by the robot. Therefore, in order to capture the voice of the speaker, the present application obtains environmental pictures collected by the robot; performs face recognition on the environmental pictures; when a face is recognized in at least one of the environmental pictures, determines the position of the face; and controls the rotation of the robot based on the position of the face so that the face is within the narrow-beam sound receiving range of the robot to wake up the robot. In this way, the robot is controlled to rotate based on the position of the face so that the recognized face is within the narrow-beam sound receiving range of the robot, thereby accurately waking up the robot. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.

[0043] Figure 1 1 is a flow chart of a robot awakening method according to an embodiment;

[0044] Figure 2 is a flow chart of a robot awakening method in another embodiment;

[0045] Figure 3 A flowchart of a tracking step in one embodiment;

[0046] Figure 4 Schematic diagram of a flow chart of a response information generating step in one embodiment;

[0047] Figure 5 A flowchart of the steps for obtaining scene description information in one embodiment;

[0048] Figure 6 is a structural block diagram of a robot awakening device in one embodiment;

[0049] Figure 7 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0050] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0051] In one embodiment, Figure 1 As shown, a robot wake-up generation method is provided. This embodiment uses the method applied to a robot terminal as an example for illustration. It is understandable that the method can also be applied to a server corresponding to the robot terminal, and can also be applied to a system including a robot terminal and a server, and is implemented through the interaction between the robot terminal and the server. In this embodiment, the method includes the following steps:

[0052] S102: Obtain the environment image collected by the robot.

[0053] The environmental image is collected by the camera on the robot, and the robot may include at least one camera. The field of view corresponding to the camera is the field of view of the robot. In order to accurately wake up the robot and avoid waking up the current robot with instructions to wake up other robots, the robot will only be woken up when the wake-up person is within the field of view of the robot.

[0054] In other embodiments, the robot includes at least two cameras, and the fields of view of the at least two cameras constitute the robot's field of view. It should be noted that the field of view of the at least two cameras can be 360 degrees or a wide-angle area, without specific limitation herein. Images are captured within this field of view to obtain an environmental image. Furthermore, when the robot includes at least two cameras, the environmental images captured by the two cameras must be timestamped to ensure that subsequently processed environmental images are images of the environment at the same time or within the same time range.

[0055] S104: Perform face recognition on the environment image.

[0056] Face recognition can be performed through a neural network model or through template matching, which is not specifically limited here.

[0057] In some optional embodiments, face recognition in an environment image may only identify a complete face. If the environment image only includes a partial face, the face is not considered recognized. A face is only considered recognized if a complete face is recognized in the environment image.

[0058] In some optional embodiments, the result of face recognition may include only one face, or at least two faces, which is not specifically limited here.

[0059] In some optional embodiments, when the environment pictures include at least two pictures, face recognition can be performed on each of the environment pictures in parallel to improve processing efficiency, and the face recognized in each environment picture can be obtained.

[0060] In some optional embodiments, if a human face is recognized, the face can also be tracked. If the time the face stays within the robot's visual range is less than a time threshold, the person corresponding to the face is considered not to be the wake-up person. The dwell time can be calculated based on the number of environmental images collected. For example, subsequent processing will only proceed if the face is included in a preset number of consecutive environmental images. The preset number of images can be 10, etc., and is not specifically limited here. The preset number of images corresponds to the time threshold, and the setting of the time threshold can be set based on the application environment of the robot and is not specifically limited here.

[0061] S106: When a face is recognized in at least one of the environment images, the position of the face is determined.

[0062] Among them, the face recognition process can not only determine whether there is a face, but also determine the position of the face in the image. Based on the intrinsic and extrinsic parameters of the camera, the position of the face in the world coordinate system can be determined.

[0063] S108: Control the robot to rotate based on the position of the human face so that the human face is within the narrow beam receiving range of the robot to wake up the robot.

[0064] A narrow beam refers to the phenomenon in which the beamwidth of electromagnetic waves (such as radio waves or laser beams) directed toward a target is narrow. Beaming refers to the focusing of electromagnetic wave energy within a specific range, forming a smaller area. In the communications and radar fields, narrow beams are used to achieve high-precision target positioning and data transmission. By focusing the beam toward the target, the reception efficiency of the target signal can be improved, the impact of interference and noise can be reduced, and the performance of communications and radar systems can be enhanced. Narrow beams can be achieved by using specialized antennas (such as high-gain antennas or array antennas), appropriate beamforming algorithms, and adjusting transmit power and receive gain. Narrow beam technology is widely used in wireless communications, satellite communications, and radar target acquisition.

[0065] In this application, only the voices of people speaking within the narrow-beam sound pickup range will be captured by the robot, and the voices of people speaking outside the narrow-beam sound pickup range will not be captured by the robot. Therefore, in order to capture the speaker's voice, in this application, the robot is controlled to rotate based on the position of the face so that the recognized face is within the robot's narrow-beam sound pickup range to successfully wake the robot. In order to ensure the strength of the acquired audio signal, the recognized face is located in the center of the robot's narrow-beam sound pickup range, so that the strength of the acquired audio signal can be maximized.

[0066] In the above-mentioned robot wake-up method, only the voices of people speaking within the narrow-beam sound receiving range will be captured by the robot, and the voices of people speaking outside the narrow-beam sound receiving range will not be captured by the robot. Therefore, in order to capture the voices of the speakers, the present application obtains environmental pictures collected by the robot; performs face recognition on the environmental pictures; when a face is recognized in at least one of the environmental pictures, determines the position of the face; and controls the rotation of the robot based on the position of the face so that the face is within the narrow-beam sound receiving range of the robot to wake up the robot. In this way, the rotation of the robot is controlled based on the position of the face so that the recognized face is within the narrow-beam sound receiving range of the robot, thereby accurately waking up the robot.

[0067] In one optional embodiment, determining the position of a face includes: in the case of identifying at least one face, determining the position of the face in parallel; the method also includes: sorting the determined positions of the face based on the relative distance between the robot's narrow beam receiving range and the position of the face.

[0068] When at least one face is identified, the position of the face can be determined in parallel. The determined face positions can then be sorted based on the relative distance between the robot's narrowbeam sound reception range and the face's position, facilitating subsequent control of the robot's rotation to determine the location of the person being woken. For example, the positions can be sorted from small to large based on the relative distance between the robot's narrowbeam sound reception range and the face's position. The faces are then positioned within the robot's narrowbeam sound reception range, and a determination is made as to whether audio is captured. If audio is not captured, the robot's rotation is continued until audio is captured. The face within the narrowbeam sound reception range corresponding to the audio is then identified as the face of the person being woken, thereby waking the robot.

[0069] In the above embodiment, since the robot's visual range is larger than the narrow beam receiving range, face recognition and tracking are first performed through vision, and the robot is controlled to rotate so that the face is within the robot's narrow beam receiving range, so that the speaker's audio can be accurately obtained to wake up the robot.

[0070] In one of the optional embodiments, combined with Figure 2 As shown, Figure 2 This is a flowchart of a robot wake-up method in another embodiment, which, after determining the position of a human face, further includes: segmenting a human face image from an environment image based on the position of the human face; performing mouth opening and closing detection on the human face image; if the human face image has an open or closed mouth, continuing to execute the step of controlling the robot rotation based on the position of the human face; if the human face image does not have an open or closed mouth, ignoring the position of the human face.

[0071] Among them, in order to ensure that the face detection has a sound source, this application also performs open and closed mouth detection on the detected face to narrow the range. Open and closed mouth detection can accurately find the speaker when there are multiple people present.

[0072] After identifying the face's location, the face image is segmented from the surrounding image. For example, the identified face can be segmented as the foreground. Then, mouth-opening and closed mouth detection is performed only on the segmented face image. Mouth-opening and closed mouth detection can also be performed based on neural networks or template matching, which are not specifically limited here.

[0073] If a face image shows an open or closed mouth, the system determines that the person in the image may be speaking. This is to avoid false triggering due to mouth opening or closing, such as coughing. The system only determines that the person in the image may be speaking, and then uses a narrow beam to pick up the sound within the range to accurately determine the person being woken up.

[0074] In the case that there is no open or closed mouth in the face image, the face corresponding to the face image is not the face of the person who wakes up. In this way, part of the face is removed through the open or closed mouth detection. For the remaining face positions, the determined face positions can be sorted based on the relative distance between the robot's narrow beam sound receiving range and the position of the face, so as to facilitate the subsequent control of the robot to rotate and determine the position of the person who wakes up. For example, based on the relative distance between the robot's narrow beam sound receiving range and the position of the face, the positions are sorted from small to large, and then the faces are placed in the robot's narrow beam sound receiving range in turn, and it is determined whether the audio is obtained. If the audio is not obtained, the robot continues to be controlled to rotate until the audio is obtained, and the face in the narrow beam sound receiving range corresponding to the audio is used as the face of the person who wakes up, thereby realizing the awakening of the robot.

[0075] In the above embodiment, after detecting the face, the mouth opening and closing detection is also performed, which can narrow the face range in the case of multiple people and improve the accuracy of the awakened person recognition.

[0076] In one of the optional embodiments, the present application further includes a tracking step, combined with Figure 3 As shown, Figure 3 The present invention is a flowchart of a tracking step in an embodiment; after controlling the robot to rotate based on the position of the face, the tracking step includes: obtaining in real time the position of the wake-up person whose face is within the narrow beam receiving range of the robot; when the position of the wake-up person changes, obtaining a first relative posture change of the wake-up person; determining a second relative posture change of the robot based on the first relative posture change; and controlling the robot to track the wake-up person based on the second relative posture change of the robot.

[0077] In order to ensure the acquisition strength and accuracy of the audio signal, the present application also includes tracking of the wake-up person. Specifically, the position of the wake-up person determined is obtained in real time. The position of the wake-up person can be determined based on vision, such as obtaining an environmental image, determining the position of the wake-up person based on the environmental image, and determining whether the position of the wake-up person has changed based on the position of the wake-up person corresponding to the adjacent environmental image (adjacent here can be understood as the previous and next frames, or the environmental image corresponding to the position acquisition period). If the position of the wake-up person has changed, the first relative posture change of the wake-up person is determined based on the adjacent environmental image. The first relative posture change is in the world coordinate system, and then the first relative posture change is processed based on the internal and external parameters of the camera to obtain the second relative posture change of the robot. The second relative posture change can only include rotation, that is, the position of the robot does not change.

[0078] In the above embodiment, when the wake-up person is determined, the wake-up person can be tracked to ensure the strength of the voice signal.

[0079] In one of the optional embodiments, the robot includes at least two cameras, and the field of view of the at least two cameras constitutes the field of view of the robot; performing face recognition on the environmental image, including: performing parallel face recognition on the environmental images captured by each camera, and determining the environmental image in which the human face exists; controlling the rotation of the robot based on the position of the human face so that the human face is within the narrow beam receiving range of the robot to wake up the robot, including: controlling the rotation of the robot in sequence based on the position of the human face so that the human face is within the narrow beam receiving range of the robot to wake up the robot, and determining the wake-up person based on the acquired audio signal.

[0080] In the case where the environment image includes at least two images, face recognition can be performed on each of the images in parallel to improve processing efficiency. The face identified in each environment image is obtained. In some cases, there may be no face in the environment image, one face in the environment image, or multiple faces in the environment image. No further processing is performed on the environment image without a face. For environment images with one or more faces, the face's location is determined. The robot's rotation is then controlled based on the face's location to ensure that the speaking face is within the robot's narrow beam reception range.

[0081] In one optional embodiment, after controlling the robot to rotate based on the position of the human face, the method includes obtaining an audio signal received by the robot; converting the audio signal into text information; calling a large model to process the text information to obtain response information corresponding to the audio signal; and outputting the response information.

[0082] After the robot is awakened, it can obtain the audio signal of a person within its narrow beam receiving range and perform subsequent processing, including converting the audio signal into text information, calling a large model to process the text information to obtain response information corresponding to the audio signal, and the response information includes voice or action, etc., which is not specifically limited here.

[0083] Among them combined Figure 4 As shown, the step of calling the large model to process the text information to obtain response information corresponding to the audio signal may include:

[0084] S402: Obtain scene description information corresponding to the current scene, where the scene description information is obtained by performing attention mechanism processing on the target scene image.

[0085] The scene description information is obtained by performing attention mechanism processing based on the target scene image, and the scene description information is used to describe the attention processing results between any two tokens in the scene. The target scene image is an image of the scene in which the robot is located, and the robot performs image acquisition in real time during operation. In some optional embodiments, the robot may include multiple cameras, each of which can capture a panoramic image of the scene in which the robot is located. In other embodiments, the robot may include only one camera or a binocular camera, and the field of view of the camera is the field of view of the robot, so that the image captured by the camera can represent the image of the scene in which the robot is located.

[0086] The attention mechanism is the attention mechanism of the decoder in the pre-trained model VLM, which is used to calculate the attention processing results between any two tokens output by the feature extraction network.

[0087] S404: Receive inquiry information.

[0088] The inquiry information is information input by the user, for example, the inquiry information may be “what is included in the scene”, “what are the details of an object in the scene”, etc. The user can make inquiries based on needs.

[0089] After receiving the query information, the query information is processed to obtain a text sequence, for example, the text is segmented to obtain a vector corresponding to each segmented word, and then the vectors are connected to obtain a text sequence corresponding to the query information. In other embodiments, other processing methods can also be used.

[0090] In addition, it should be noted that if the inquiry information is in the form of voice, in this application, the inquiry information is first converted from voice to text to obtain text information, and then the text information is processed.

[0091] S406: Input the query information and the cached scene description information into the pre-trained model to obtain a response result corresponding to the query information.

[0092] The pre-trained model VLM includes an attention mechanism, which needs to calculate the attention results between any two tokens. In this application, it is assumed that the scene description information includes the attention results between three tokens, namely A, B and C. The scene description information includes the attention results between AB, the attention results between AC and the attention results between BC. Now there is query information. If you want to get the corresponding image-text conversion results, you also need to calculate the attention results between the query information Q and the above three tokens, that is, the attention results between AQ, the attention results between BQ and the attention results between CQ.

[0093] In traditional technology, each time a query is received, the scene image and query need to be input into the VLM. First, feature extraction is performed on the scene image, and text extraction is performed on the query. Then, the attention mechanism is calculated through transformation. Taking the above example, it is necessary to calculate the attention results between AB, the attention results between AC, the attention results between AQ, the attention results between BC, the attention results between BQ, and the attention results between CQ. These attention results need to be calculated each time a query is input, resulting in a large amount of repeated calculations.

[0094] To address the aforementioned technical issues, this application stores the attention results between AB, AC, and BC as scene description information. This allows subsequent processing to calculate only the attention results between each token of the scene image and the query information Q. In this application, each token of the scene image is also cached. This reduces the repeated calculation of scene description information involved in the query processing, improving processing efficiency.

[0095] The above-mentioned robot question-answer response generation method first obtains the scene image, and obtains the scene description information based on the scene image, and caches it in advance. When making subsequent inquiries, it is directly processed based on the query information and the scene description information without having to recalculate the attention results between the scene description information. This makes full use of the scene description information, reduces the data processing amount of the model, and improves processing efficiency.

[0096] In one of the optional embodiments, combined with Figure 5 As shown, Figure 5This is a flowchart of a scene description information acquisition step in an embodiment. The scene description information acquisition step, i.e., acquiring scene description information corresponding to the current scene, includes: acquiring a target scene image; inputting the target scene image into a pre-trained model, and using the processing result of the attention mechanism in the pre-trained model as scene description information; and caching the scene description information.

[0097] The pre-trained model VLM is a model that includes robot vision and a large prediction model. The structure of the large prediction model can be a transform structure, so that the transform's attention mechanism can be used, such as the cross-attention mechanism, to calculate the cross-attention results of each element in the scene.

[0098] The pre-trained model VLM includes a feature extraction network and a decoder corresponding to the transform. The feature extraction network is used to extract the scene sequence corresponding to the scene image, and then input the scene sequence into the transform for decoding. The decoding process of the transform includes the processing of the cross-attention mechanism, which can calculate the attention results between different elements in the scene sequence and cache the attention results as scene description information.

[0099] In the above embodiment, scene description information is generated in each new scene, so when the query is subsequently processed, there is no need to process the scene image, but the scene description information and the query are directly processed, which reduces the repeated calculation of the scene description information and improves the calculation efficiency.

[0100] In one optional embodiment, obtaining the target scene image includes: determining a difference between the current scene image and the historical scene image; and using the current scene image as the target scene image when the difference is greater than a threshold.

[0101] The current scene image is a scene image captured in real time by the robot's camera, and the historical scene image is a scene image previously captured by the robot's camera. The historical scene image can be the target scene image of the previous scene, or any scene image captured from the previous scene to the current time. Optionally, the present application uses the historical scene image as the target scene image of the previous scene as an example for explanation.

[0102] Calculating the difference between the current scene image and the historical scene image can be performed based on SLAM. Specifically, calculating the difference between the current scene image and the historical scene image can be performed by calculating the difference in feature points between the current scene image and the historical scene image. For example, the common parts in the current scene image and the historical scene image can be calculated, and then the ratio of the number of pixels in the common parts to the number of pixels in the current scene image can be used as the difference. In other embodiments, the difference can also be determined based on optical flow information, fixed objects, etc., which are not specifically limited here.

[0103] The threshold is preset and can be obtained based on experience, such as 75%. In other embodiments, other values can also be selected and are not specifically limited here.

[0104] When the difference is greater than the threshold, it is determined that the scene has changed, and it is necessary to obtain the target scene image corresponding to the scene. Optionally, in this application, for ease of processing, the current scene image with a difference greater than the threshold is directly used as the target scene image, that is, the first frame scene image entering the new scene.

[0105] In the above embodiment, the scene change can be judged based on the context scene image, thereby ensuring the real-time performance of the target scene image and further ensuring the accuracy of the scene description information.

[0106] In one of the optional embodiments, the target scene image is input into a pre-trained model, and the processing result of the attention mechanism in the pre-trained model is used as the scene description information, including: extracting the scene sequence corresponding to the scene image through the feature extraction network of the pre-trained model; using the scene sequence as the input of the decoder, and using the attention mechanism of the decoder to calculate the attention results between different elements in the scene sequence; and using the attention results as the scene description information.

[0107] Among them, the feature extraction network can be a simple neural network, which can extract the scene sequence corresponding to the scene image. Optionally, the feature extraction network can be connected to a fully connected layer. After the scene sequence is extracted by the scene extraction network, it is then classified through the fully connected layer to complete the prediction of each token, and then the obtained token is input into the decoder as a scene sequence.

[0108] The decoder is transform, which may include an attention mechanism. The attention mechanism may include a cross-attention mechanism, through which the attention result between any two tokens can be calculated.

[0109] In the above embodiment, after entering a new scene, the scene description information is first calculated as the basis for subsequent inquiries. In this way, there is no need to calculate the attention results between any two tokens in the scene during subsequent inquiries, which can improve processing efficiency.

[0110] In one of the optional embodiments, the scene description information is cached, including: storing the scene description information in a k-value cache, wherein the index k of the scene description information in the k-value cache is a token, and the value value is the attention result between any two tokens.

[0111] The scene description information includes the attention results between different tokens, where the index can be a token and the value is the attention result between any two tokens.

[0112] In one of the optional embodiments, combined with Figure 3 As shown, Figure 3 This is a processing flow chart for an embodiment in which a response result is that no answer corresponding to the query information is obtained. The query information and cached scene description information are input into a pre-trained model to obtain a response result corresponding to the query information, including: when the response result is that no answer corresponding to the query information is obtained, keyword extraction is performed on the query information; image segmentation is performed on the target scene image based on the extracted keywords to obtain a region of interest; and the query information and the region of interest are input into the pre-trained model for processing to obtain a target response result.

[0113] Among them, the response result may include the answer corresponding to the inquiry information, and the answer corresponding to the inquiry information that cannot be obtained. If the response result is the answer corresponding to the inquiry information, the response result can be directly output. If the response result is that the answer corresponding to the inquiry information cannot be obtained, one way is to output the answer that cannot be obtained. Another way is to generate response information in combination with context information when the response result is that the answer corresponding to the inquiry information cannot be obtained.

[0114] The keyword extraction of the query information can be performed through a large prediction model. After obtaining the target scene image, the target scene image is processed through the VLM, and the scene elements, that is, the element information corresponding to the token, such as table, chair, etc. can also be output.

[0115] By using the big prediction model to extract keywords from query information, keywords can be extracted from the query based on the scene elements output by the VLM, such as the keyword "table".

[0116] Image segmentation of the target scene image based on the extracted keywords can be performed by inputting the keywords and the target scene image into a segmentation model. The segmentation model can be a SAM model, etc., which is not specifically limited here. The target scene image is segmented based on the keywords by the segmentation model to determine the region of interest corresponding to the query information, and then the region of interest is segmented from the target scene image. This can reduce the amount of image processing.

[0117] Finally, the region of interest and query information are input into the VLM for processing, where the feature extraction network of the VLM is used to extract the image sequence of the region of interest, and the text processing module is used to extract the text features of the query information to obtain a text sequence. The image sequence and text sequence are then input into the VLM decoder, and the attention mechanism of the VLM decoder is used to perform attention processing on the image sequence and text sequence, and finally the target response result is obtained.

[0118] In the above embodiment, when no answer corresponding to the query information is obtained, the contextual query information can be used to segment the scene image to obtain the region of interest, and then VLM processing can be performed based on the query information and the scene image to obtain the target response result, which makes full use of the contextual information and makes the robot response result more accurate.

[0119] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0120] Based on the same inventive concept, embodiments of the present application also provide a robot wake-up device for implementing the aforementioned robot wake-up method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more robot wake-up device embodiments provided below can be found in the above-described limitations of the robot wake-up method and will not be further elaborated here.

[0121] In an exemplary embodiment, Figure 6As shown, a robot awakening device is provided, comprising: an environment image acquisition module 601, a face recognition module 602, a face position determination module 603 and a control module 604, wherein:

[0122] The environment image acquisition module 601 is used to acquire the environment image collected by the robot;

[0123] A face recognition module 602 is used to perform face recognition on the environment image;

[0124] A face position determination module 603 is configured to determine the position of a face when a face is recognized in at least one of the environment images;

[0125] The control module 604 is used to control the rotation of the robot based on the position of the human face so that the human face is located within the narrow beam receiving range of the robot to wake up the robot.

[0126] In one optional embodiment, the face position determination module 603 is specifically configured to determine the position of the face in parallel when at least one face is identified.

[0127] The above-mentioned device also includes: a sorting module, which is used to sort the determined positions of the human face based on the relative distance between the narrow beam receiving range of the robot and the position of the human face.

[0128] In one of the optional embodiments, the above-mentioned device also includes: an open and closed mouth detection module, which is used to segment a face image from an environmental image based on the position of the face; perform open and closed mouth detection on the face image; if the face image has an open or closed mouth, continue to execute the step of controlling the rotation of the robot based on the position of the face; if the face image does not have an open or closed mouth, ignore the position of the face.

[0129] In one of the optional embodiments, the above-mentioned device also includes: a tracking module, which is used to obtain in real time the position of the wake-up person whose face is located within the narrow beam receiving range of the robot; when the position of the wake-up person changes, obtain the first relative posture change of the wake-up person; determine the second relative posture change of the robot based on the first relative posture change; and control the robot to track the wake-up person based on the second relative posture change of the robot.

[0130] In one of the optional embodiments, the robot includes at least two cameras, and the field of view of at least two cameras constitutes the field of view of the robot; the above-mentioned face recognition module 602 is specifically used to perform parallel face recognition on the environmental images captured by each camera, and determine the environmental images in which human faces exist.

[0131] The control module 604 is specifically configured to control the rotation of the robot based on the position of the human face so that the human face is within the narrow beam receiving range of the robot to wake up the robot, and to determine the wake-up person based on the acquired audio signal.

[0132] In one of the optional embodiments, the above-mentioned device also includes: a response module, used to obtain the audio signal received by the robot; convert the audio signal into text information; call the large model to process the text information to obtain response information corresponding to the audio signal; and output the response information.

[0133] Each module in the robot wake-up device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.

[0134] In an exemplary embodiment, a computer device is provided. The computer device may be a robot terminal, and its internal structure diagram may be as shown in FIG. Figure 7 As shown. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless means, and the wireless means can be implemented via Wi-Fi, a mobile cellular network, near-field communication (NFC), or other technologies. When executed by the processor, the computer program implements a robot wake-up method. The display unit of the computer device is used to form a visually visible image, and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.

[0135] Those skilled in the art will understand that Figure 7The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0136] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented: obtaining environmental images collected by a robot; performing facial recognition on the environmental images; determining the position of the face when a face is recognized in at least one of the environmental images; and controlling the rotation of the robot based on the position of the face so that the face is within the narrow beam receiving range of the robot to wake up the robot.

[0137] In one embodiment, the determination of the position of a human face implemented by the processor when executing a computer program includes: in the case of identifying at least one human face, determining the position of the human face in parallel; when the processor executes the computer program, the following steps are also implemented: based on the relative distance between the robot's narrow beam receiving range and the position of the human face, sorting the determined positions of the human face.

[0138] In one embodiment, after determining the position of a human face when the processor executes a computer program, the method further includes: segmenting a human face image from an environmental image based on the position of the human face; performing mouth opening and closing detection on the human face image; if the human face image has an open or closed mouth, continuing to execute the step of controlling the rotation of the robot based on the position of the human face; if the human face image does not have an open or closed mouth, ignoring the position of the human face.

[0139] In one embodiment, after the robot rotation is controlled based on the position of the human face when the processor executes the computer program, it includes: obtaining the position of the wake-up person whose face is located in the narrow beam receiving range of the robot in real time; when the position of the wake-up person changes, obtaining the first relative posture change of the wake-up person; determining the second relative posture change of the robot based on the first relative posture change; and controlling the robot to track the wake-up person based on the second relative posture change of the robot.

[0140] In one embodiment, the cameras on the robot involved when the processor executes the computer program include at least two, and the field of view of the at least two cameras constitutes the field of view of the robot; the face recognition of the environmental image implemented by the processor when executing the computer program includes: parallel face recognition of the environmental images captured by each camera, and determining the environmental image containing the human face; the robot rotation control based on the position of the face implemented when the processor executes the computer program so that the human face is within the narrow beam receiving range of the robot to wake up the robot includes: controlling the robot rotation in sequence based on the position of the face so that the human face is within the narrow beam receiving range of the robot to wake up the robot, and determining the wake-up person based on the acquired audio signal.

[0141] In one embodiment, after the processor executes the computer program, the position-based control of the robot rotation based on the human face includes: obtaining the audio signal received by the robot; converting the audio signal into text information; calling the large model to process the text information to obtain response information corresponding to the audio signal; and outputting the response information.

[0142] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: obtaining environmental images collected by a robot; performing facial recognition on the environmental images; determining the position of the face when a face is recognized in at least one of the environmental images; and controlling the rotation of the robot based on the position of the face so that the face is within the narrow beam receiving range of the robot to wake up the robot.

[0143] In one embodiment, the computer program implemented by the processor for determining the position of a human face includes: determining the position of the human face in parallel when at least one human face is identified; and the computer program implemented by the processor for the processor further implements the following steps: sorting the determined positions of the human faces based on the relative distance between the robot's narrow beam receiving range and the position of the human face.

[0144] In one embodiment, after determining the position of a human face when the computer program is executed by the processor, the method further includes: segmenting a human face image from an environmental image based on the position of the human face; performing mouth opening and closing detection on the human face image; if the human face image has an open or closed mouth, continuing to execute the step of controlling the rotation of the robot based on the position of the human face; if the human face image does not have an open or closed mouth, ignoring the position of the human face.

[0145] In one embodiment, after the computer program is executed by the processor, the robot rotation is controlled based on the position of the human face, including: obtaining the position of the wake-up person whose face is located within the narrow beam receiving range of the robot in real time; when the position of the wake-up person changes, obtaining the first relative posture change of the wake-up person; determining the second relative posture change of the robot based on the first relative posture change; and controlling the robot to track the wake-up person based on the second relative posture change of the robot.

[0146] In one embodiment, the cameras on the robot involved when the computer program is executed by the processor include at least two, and the field of view of the at least two cameras constitutes the field of view of the robot; the face recognition of the environmental image implemented when the computer program is executed by the processor includes: parallel face recognition of the environmental image captured by each camera, and determining the environmental image in which the human face exists; the robot rotation control based on the position of the face implemented when the computer program is executed by the processor so that the human face is within the narrow beam receiving range of the robot to wake up the robot includes: controlling the robot rotation in sequence based on the position of the face so that the human face is within the narrow beam receiving range of the robot to wake up the robot, and determining the wake-up person based on the acquired audio signal.

[0147] In one embodiment, after the computer program is executed by a processor, the robot rotation is controlled based on the position of the human face, including: obtaining an audio signal received by the robot; converting the audio signal into text information; calling a large model to process the text information to obtain response information corresponding to the audio signal; and outputting the response information.

[0148] In one embodiment, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the following steps: obtaining environmental images collected by a robot; performing facial recognition on the environmental images; determining the position of the face when a face is recognized in at least one of the environmental images; and controlling the rotation of the robot based on the position of the face so that the face is within the narrow beam reception range of the robot to wake up the robot.

[0149] In one embodiment, the computer program implemented by the processor for determining the position of a human face includes: determining the position of the human face in parallel when at least one human face is identified; and the computer program implemented by the processor for the processor further implements the following steps: sorting the determined positions of the human faces based on the relative distance between the robot's narrow beam receiving range and the position of the human face.

[0150] In one embodiment, after determining the position of a human face when the computer program is executed by the processor, the method further includes: segmenting a human face image from an environmental image based on the position of the human face; performing mouth opening and closing detection on the human face image; if the human face image has an open or closed mouth, continuing to execute the step of controlling the rotation of the robot based on the position of the human face; if the human face image does not have an open or closed mouth, ignoring the position of the human face.

[0151] In one embodiment, after the computer program is executed by the processor, the robot rotation is controlled based on the position of the human face, including: obtaining the position of the wake-up person whose face is located within the narrow beam receiving range of the robot in real time; when the position of the wake-up person changes, obtaining the first relative posture change of the wake-up person; determining the second relative posture change of the robot based on the first relative posture change; and controlling the robot to track the wake-up person based on the second relative posture change of the robot.

[0152] In one embodiment, the cameras on the robot involved when the computer program is executed by the processor include at least two, and the field of view of the at least two cameras constitutes the field of view of the robot; the face recognition of the environmental image implemented when the computer program is executed by the processor includes: parallel face recognition of the environmental image captured by each camera, and determining the environmental image in which the human face exists; the robot rotation control based on the position of the face implemented when the computer program is executed by the processor so that the human face is within the narrow beam receiving range of the robot to wake up the robot includes: controlling the robot rotation in sequence based on the position of the face so that the human face is within the narrow beam receiving range of the robot to wake up the robot, and determining the wake-up person based on the acquired audio signal.

[0153] In one embodiment, after the computer program is executed by a processor, the robot rotation is controlled based on the position of the human face, including: obtaining an audio signal received by the robot; converting the audio signal into text information; calling a large model to process the text information to obtain response information corresponding to the audio signal; and outputting the response information.

[0154] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0155] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.

[0156] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0157] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A robot awakening method, characterized in that: The method comprises: Get the environment pictures collected by the robot; Performing face recognition on the environment picture; When a human face is recognized in at least one of the environment images, determining a position of the human face; Controlling the robot to rotate based on the position of the human face so that the human face is within a narrow beam sound receiving range of the robot to wake up the robot, wherein the robot's visual range is larger than the narrow beam sound receiving range; After controlling the robot to rotate based on the position of the human face, the method further comprises: Get the audio signal received by the robot; Converting the audio signal into text information; Calling the large model to process the text information to obtain response information corresponding to the audio signal, including: obtaining scene description information corresponding to the current scene, the scene description information is obtained by performing attention mechanism processing based on the target scene image, inputting the text information and the cached scene description information into the pre-trained model to obtain a response result corresponding to the text information, wherein the scene description information is an attention processing result between any two elements obtained by performing attention mechanism processing based on the target scene image; The response information is output.

2. The method according to claim 1, characterized in that Determining the position of the face includes: When at least one face is identified, the position of the face is determined in parallel; Before controlling the rotation of the robot based on the position of the human face, the method further includes: The determined positions of the human faces are sorted based on the relative distance between the narrow beam receiving range of the robot and the positions of the human faces.

3. The method according to claim 2, characterized in that After determining the position of the face, the method further includes: Segmenting the face image from the environment image based on the position of the face; Performing mouth opening and closing detection on the face image; If the face image contains an open or closed mouth, continuing to execute the step of controlling the rotation of the robot based on the position of the face; When the face image does not have an open or closed mouth, the position of the face is ignored.

4. The method according to any one of claims 1 to 3, characterized in that After controlling the robot to rotate based on the position of the human face, the method further comprises: Real-time acquisition of the wake-up person's location where the face is within the robot's narrow beam reception range; When the position of the awakening person changes, obtaining a first relative posture change of the awakening person; determining a second relative posture change of the robot based on the first relative posture change; The robot is controlled to track and wake up the person based on the change of the second relative posture of the robot.

5. The method according to any one of claims 1 to 3, characterized in that The robot includes at least two cameras, and the field of view of the at least two cameras constitutes the field of view of the robot; performing face recognition on the environment image includes: Performing parallel face recognition on the environmental images captured by the cameras, and determining environmental images containing human faces; The controlling the rotation of the robot based on the position of the human face so that the human face is located within the narrow beam receiving range of the robot to wake up the robot includes: The robot is controlled to rotate in sequence based on the position of the human face so that the human face is located within the narrow beam sound receiving range of the robot to wake up the robot, and the awakened person is determined based on the acquired audio signal.

6. A robot awakening device, characterized in that: The device comprises: Environmental image acquisition module, used to obtain environmental images collected by the robot; A face recognition module, used to perform face recognition on the environment image; a face position determination module, configured to determine the position of a face when a face is recognized in at least one of the environment images; a control module, configured to control the rotation of the robot based on the position of the human face so that the human face is within the narrow beam receiving range of the robot to wake up the robot, wherein the robot's visual range is larger than the narrow beam receiving range; The device further comprises: The response module is used to obtain the audio signal received by the robot; convert the audio signal into text information; call the large model to process the text information to obtain response information corresponding to the audio signal, including: obtaining scene description information corresponding to the current scene, the scene description information is obtained based on the attention mechanism processing of the target scene image, inputting the text information and the cached scene description information into the pre-trained model to obtain a response result corresponding to the text information, wherein the scene description information is the attention processing result between any two elements obtained based on the attention mechanism processing of the target scene image; and outputting the response information.

7. The device according to claim 6, characterized in that The face position determination module is specifically configured to determine the position of the face in parallel when at least one face is identified; The device further comprises: The sorting module is used to sort the determined positions of the human faces based on the relative distance between the narrow beam receiving range of the robot and the position of the human face.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Sound pickup method and device

    CN109688512A

  • Voice interaction wake-up-free method and device

    CN112634895A

  • Control method and device of vehicle-mounted intelligent equipment, vehicle and storage medium

    CN116080565A

  • Virtual digital human interaction device, system and method

    CN118535005A

  • Voice interactive robot and voice interaction system

    US20180311816A1