Robot awakening method and device, computer equipment, readable storage medium and program product

By identifying the position of the face in the environmental picture and controlling the rotation of the robot, the face is within the narrow beam radio range, the problem of waste of resources and difficulty in memory of wake-up words in traditional awakening methods is solved, and the accuracy and efficiency of the robot awakening are achieved.

CN119992631AActive Publication Date: 2025-05-13SHANGHAI FOURIER INTELLIGENCE CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510473556.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

Traditional robot awakening methods require specific wake-up words, which leads to waste of resources and difficulty in memory of wake-up words, making it difficult to accurately awaken the robot.

Method used

By obtaining the environment pictures collected by the robot, facial recognition is performed, the face position is determined, and the robot rotation is controlled based on the face position, so that the face is within the narrow beam sound range of the robot, thereby accurately waking the robot.

Benefits of technology

The accuracy and efficiency of robot awakening are achieved, and the difficulty of waste of resources and awakening word memory is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992631A_ABST
    Figure CN119992631A_ABST
Patent Text Reader

Abstract

The invention relates to a robot awakening method and device, computer equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: acquiring an environment picture collected by a robot; performing face recognition on the environment picture; under the condition that at least one environment picture is recognized to obtain a human face, determining the position of the human face; and controlling the robot to rotate based on the position of the human face, so that the human face is located in a narrow beam reception range of the robot to wake up the robot. By adopting the method, the robot can be accurately awakened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of robot control technology, and in particular to a robot awakening method, device, computer equipment, computer-readable storage medium and computer program product. Background Art

[0002] With the rapid development of modern technology, the application scope of intelligent robots is becoming more and more extensive. Intelligent robots can be seen in homes, shopping malls, banks and other public places.

[0003] However, in traditional technology, a specific wake-up word is required to wake up a robot. Different types of robots have different wake-up words, while robots of the same type have the same wake-up words. In this way, when the same robot is used, many robots can be awakened by the same wake-up word, resulting in a waste of robot resources. In addition, if different types of robots are used, each robot has a different wake-up word, and it is also challenging to remember these wake-up words. Therefore, there is an urgent need for a way to accurately wake up the robot. Summary of the invention

[0004] Based on this, it is necessary to provide a robot wake-up method, device, computer equipment, computer-readable storage medium and computer program product that can accurately wake up the robot in response to the above technical problems.

[0005] In a first aspect, the present application provides a robot awakening method, the method comprising:

[0006] Get the environment pictures collected by the robot;

[0007] Performing face recognition on the environment picture;

[0008] When a face is recognized in at least one of the environment images, determining a position of the face;

[0009] The robot is controlled to rotate based on the position of the human face so that the human face is located within the narrow beam sound receiving range of the robot to wake up the robot.

[0010] In one embodiment, determining the position of the face includes:

[0011] When at least one face is identified, the position of the face is determined in parallel;

[0012] Before controlling the robot to rotate based on the position of the human face, the method further includes:

[0013] The determined positions of the human faces are sorted based on the relative distance between the narrow beam sound receiving range of the robot and the position of the human face.

[0014] In one embodiment, after determining the position of the face, the method further includes:

[0015] Segmenting the face image from the environment image based on the position of the face;

[0016] Performing open and closed mouth detection on the face image;

[0017] In the case where the face image has an open or closed mouth, continuing to execute the step of controlling the robot to rotate based on the position of the face;

[0018] When the face image does not have an open or closed mouth, the position of the face is ignored.

[0019] In one embodiment, after controlling the robot to rotate based on the position of the human face, the method further comprises:

[0020] Real-time acquisition of the location of the wake-up person whose face is within the robot's narrow beam receiving range;

[0021] When the position of the awakening person changes, acquiring a first relative posture change of the awakening person;

[0022] Determining a second relative posture change of the robot based on the first relative posture change;

[0023] Based on the change of the second relative posture of the robot, the robot is controlled to track and wake up the person.

[0024] In one embodiment, the robot includes at least two cameras, and the field of view of at least two of the cameras constitutes the field of view of the robot; performing face recognition on the environment image includes:

[0025] Performing parallel face recognition on the environmental images collected by the cameras, and determining environmental images containing human faces;

[0026] The controlling the robot to rotate based on the position of the human face so that the human face is located within the narrow beam sound receiving range of the robot to wake up the robot includes:

[0027] The robot is controlled to rotate in sequence based on the position of the human face so that the human face is located within the narrow beam sound receiving range of the robot to wake up the robot, and the awakened person is determined based on the acquired audio signal.

[0028] In one embodiment, after controlling the robot to rotate based on the position of the human face, the method further comprises:

[0029] Obtain the audio signal received by the robot;

[0030] Converting the audio signal into text information;

[0031] Calling the large model to process the text information to obtain response information corresponding to the audio signal;

[0032] The response information is output.

[0033] In a second aspect, the present application further provides a robot awakening device, the device comprising:

[0034] The environment image acquisition module is used to obtain the environment images collected by the robot;

[0035] A face recognition module, used for performing face recognition on the environment image;

[0036] A face position determination module, used to determine the position of the face when a face is recognized in at least one of the environment images;

[0037] The control module is used to control the rotation of the robot based on the position of the human face so that the human face is located within the narrow beam receiving range of the robot to wake up the robot.

[0038] In a third aspect, the present application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method in any one of the above-mentioned embodiments when executing the computer program.

[0039] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method in any one of the above-mentioned embodiments.

[0040] In a fifth aspect, the present application also provides a computer program product, including a computer program, which implements the steps of the method in any one of the above embodiments when executed by a processor.

[0041] According to the above-mentioned robot wake-up method, device, computer equipment, computer-readable storage medium and computer program product, only the voice of a person speaking within the narrow-beam sound receiving range will be captured by the robot, and the voice of a person speaking outside the narrow-beam sound receiving range will not be captured by the robot. Therefore, in order to capture the voice of the speaker, the present application obtains environmental pictures collected by the robot; performs face recognition on the environmental pictures; when a face is recognized in at least one of the environmental pictures, determines the position of the face; and controls the robot to rotate based on the position of the face so that the face is within the narrow-beam sound receiving range of the robot to wake up the robot. In this way, the robot is controlled to rotate based on the position of the face so that the recognized face is within the narrow-beam sound receiving range of the robot to accurately wake up the robot. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0043] Figure 1 A schematic diagram of a process of waking up a robot in one embodiment;

[0044] Figure 2 is a flow chart of a robot awakening method in another embodiment;

[0045] Figure 3 A flowchart of a tracking step in an embodiment;

[0046] Figure 4 A schematic diagram of a flow chart of a response information generating step in one embodiment;

[0047] Figure 5 is a flowchart of a step of acquiring scene description information in an embodiment;

[0048] Figure 6 is a structural block diagram of a robot awakening device in one embodiment;

[0049] Figure 7 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0050] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0051] In one embodiment, Figure 1 As shown, a robot wake-up generation method is provided. This embodiment takes the method applied to a robot terminal as an example. It can be understood that the method can also be applied to a server corresponding to the robot terminal, and can also be applied to a system including a robot terminal and a server, and is implemented through the interaction between the robot terminal and the server. In this embodiment, the method includes the following steps:

[0052] S102: Obtain the environment image collected by the robot.

[0053] The environmental image is collected by the camera on the robot, and the robot may include at least one camera. The field of view corresponding to the camera is the field of view of the robot. In order to accurately wake up the robot and avoid waking up the current robot with instructions to wake up other robots, the robot will only be woken up when the waker is within the field of view of the robot.

[0054] In other embodiments, the robot includes at least two cameras, and the field of view of the at least two cameras constitutes the field of view of the robot, wherein it should be noted that the field of view of the at least two cameras can be 360 ​​degrees, or a wide-angle area, which is not specifically limited here, and the image is collected within the field of view to obtain the environment picture. In addition, it should be noted that when the robot includes at least two cameras, the environment pictures collected by the two cameras need to be timestamped to ensure that the environment pictures processed subsequently are environment pictures at the same time point or within the same time range.

[0055] S104: Perform face recognition on the environment image.

[0056] The face recognition may be performed through a neural network model or through template matching, which is not specifically limited here.

[0057] In some optional embodiments, face recognition on the environment image may only identify a complete face. If the environment image only includes a partial face, the face is not considered to be recognized. Only if a complete face is recognized in the environment image, the face is considered to be recognized.

[0058] In some optional embodiments, the result of face recognition may include only one face, or at least two faces, which is not specifically limited here.

[0059] In some optional embodiments, when the environment pictures include at least two pictures, face recognition can be performed on each environment picture in parallel to improve processing efficiency, and the face recognized by each environment picture can be obtained.

[0060] In some optional embodiments, when a human face is recognized, the human face can also be tracked. If the residence time of the human face in the robot's visual range is less than the time threshold, it is considered that the person corresponding to the human face is not the wake-up person. The residence time can be calculated based on the number of environmental images collected. For example, the face is included in a preset number of consecutive environmental images before subsequent processing is performed. The preset number of images can be 10, etc., which is not specifically limited here. The preset number of images corresponds to the time threshold, and the setting of the time threshold can be set based on the application environment of the robot, which is not specifically limited here.

[0061] S106: When a face is recognized in at least one of the environment images, determine the position of the face.

[0062] Among them, the face recognition process can not only determine whether there is a face, but also determine the position of the face in the image. Based on the intrinsic and extrinsic parameters of the camera, the position of the face in the world coordinate system can be determined.

[0063] S108: Control the robot to rotate based on the position of the human face so that the human face is located within the narrow beam receiving range of the robot to wake up the robot.

[0064] Among them, narrow beam refers to the phenomenon that the beam width of electromagnetic waves (such as radio waves or laser beams) directed at the target is narrow. Beam refers to the energy of electromagnetic waves focused within a certain range to form a smaller area. In the field of communication and radar, narrow beams are used to achieve high-precision target positioning and data transmission. By focusing the beam directed to the target, the reception efficiency of the target signal can be improved, and the influence of interference and noise can be reduced, thereby improving the performance of communication and radar systems. The generation of narrow beams can be achieved by using special antennas (such as high-gain antennas or array antennas), appropriate beamforming algorithms, and adjusting the transmit power and receive gain. Narrow beam technology is widely used in wireless communications, satellite communications, radar target acquisition, etc.

[0065] In this application, only the voices of people who are within the narrow beam sound receiving range will be captured by the robot, and the voices of people who are outside the narrow beam sound receiving range will not be captured by the robot. Therefore, in order to capture the voice of the speaker, in this application, the robot is controlled to rotate based on the position of the face so that the recognized face is within the narrow beam sound receiving range of the robot to successfully wake up the robot. In order to ensure the strength of the acquired audio signal, the recognized face is located in the center of the narrow beam sound receiving range of the robot, so that the strength of the acquired audio signal can be maximized.

[0066] In the above robot wake-up method, only the voices of people speaking within the narrow-beam sound receiving range will be captured by the robot, and the voices of people speaking outside the narrow-beam sound receiving range will not be captured by the robot. Therefore, in order to capture the voice of the speaker, the present application obtains environmental pictures collected by the robot; performs face recognition on the environmental pictures; when a face is recognized in at least one of the environmental pictures, determines the position of the face; and controls the robot to rotate based on the position of the face so that the face is within the narrow-beam sound receiving range of the robot to wake up the robot. In this way, the robot is controlled to rotate based on the position of the face so that the recognized face is within the narrow-beam sound receiving range of the robot, thereby accurately waking up the robot.

[0067] In one of the optional embodiments, determining the position of a human face includes: determining the position of the human face in parallel when at least one human face is identified; the method also includes: sorting the determined positions of the human faces based on the relative distance between the robot's narrow beam receiving range and the position of the human face.

[0068] When at least one face is identified, the position of the face can be determined in parallel, and the positions of the determined faces can be sorted based on the relative distance between the robot's narrow beam sound receiving range and the position of the face, so as to facilitate the subsequent control of the robot to rotate and determine the position of the wake-up person. For example, the positions are sorted from small to large based on the relative distance between the robot's narrow beam sound receiving range and the position of the face, and then the faces are placed in the robot's narrow beam sound receiving range in turn, and it is determined whether the audio is obtained. If the audio is not obtained, the robot continues to be controlled to rotate until the audio is obtained, and the face in the narrow beam sound receiving range corresponding to the audio is used as the face of the wake-up person, thereby realizing the awakening of the robot.

[0069] In the above embodiment, since the robot's visual range is larger than the narrow beam receiving range, face recognition and tracking are first performed through vision, and the robot is controlled to rotate so that the face is within the robot's narrow beam receiving range, so that the speaker's audio can be accurately obtained to wake up the robot.

[0070] In one of the optional embodiments, in combination Figure 2 As shown, Figure 2 This is a flowchart of a robot wake-up method in another embodiment, which, after determining the position of a human face, also includes: segmenting a face image from an environment image based on the position of the face; performing open or closed mouth detection on the face image; if the face image has an open or closed mouth, continuing to execute the step of controlling the rotation of the robot based on the position of the face; if the face image does not have an open or closed mouth, ignoring the position of the face.

[0071] Among them, in order to ensure that the face detection has a sound source, the detected face is also subjected to open and closed mouth detection to narrow the range. The open and closed mouth detection can accurately find the speaker when multiple people are present.

[0072] After the position of the face is identified, the face image is segmented from the environment image. For example, the identified face can be segmented as the foreground. Then, only the segmented face image is subjected to the mouth opening and closing detection. The mouth opening and closing detection can also be performed based on a neural network or template matching, which is not specifically limited here.

[0073] If the face image has an open or closed mouth, it is determined that the face in the face image may be speaking. This is to avoid false triggering of the opening or closing of the mouth caused by coughing and other actions. Here, it is only determined that the face in the face image may be speaking, and then the sound is collected within the narrow beam receiving range to determine the correct person to wake up.

[0074] If there is no open or closed mouth in the face image, the face corresponding to the face image is not the face of the person who wakes up. In this way, part of the face is removed through the open or closed mouth detection. For the remaining face positions, the determined face positions can be sorted based on the relative distance between the robot's narrow beam sound receiving range and the face position, so as to facilitate the subsequent control of the robot to rotate and determine the position of the person who wakes up. For example, based on the relative distance between the robot's narrow beam sound receiving range and the face position, the positions are sorted from small to large, and then the face is placed in the robot's narrow beam sound receiving range in turn, and it is determined whether the audio is obtained. If the audio is not obtained, the robot continues to be controlled to rotate until the audio is obtained, and the face in the narrow beam sound receiving range corresponding to the audio is used as the face of the person who wakes up, thereby realizing the awakening of the robot.

[0075] In the above embodiment, after the face is detected, the open and closed mouth detection is also performed, which can narrow the face range in the case of multiple people and improve the accuracy of the awakened person recognition.

[0076] In one of the optional embodiments, the present application further includes a tracking step, combined with Figure 3 As shown, Figure 3 The present invention is a flowchart of a tracking step in an embodiment; after the robot is controlled to rotate based on the position of a human face, the tracking step includes: obtaining in real time the position of a wake-up person whose face is within the narrow beam receiving range of the robot; when the position of the wake-up person changes, obtaining a first relative posture change of the wake-up person; determining a second relative posture change of the robot based on the first relative posture change; and controlling the robot to track the wake-up person based on the second relative posture change of the robot.

[0077] Among them, in order to ensure the acquisition strength and accuracy of the audio signal, the present application also includes tracking of the wake-up person, specifically, obtaining the determined position of the wake-up person in real time, and the position of the wake-up person can be determined based on vision, such as obtaining an environmental picture, determining the position of the wake-up person based on the environmental picture, and determining whether the position of the wake-up person has changed based on the position of the wake-up person corresponding to the adjacent environmental picture (the adjacent here can be understood as the previous and next frames, and can also be understood as the environmental picture corresponding to the position acquisition cycle). If the position of the wake-up person has changed, the first relative posture change of the wake-up person is determined based on the adjacent environmental picture. The first relative posture change is in the world coordinate system, and then the first relative posture change is processed based on the internal and external parameters of the camera to obtain the second relative posture change of the robot. The second relative posture change can only include rotation, that is, the position of the robot does not change.

[0078] In the above embodiment, when the wake-up person is determined, the wake-up person can be tracked to ensure the strength of the voice signal.

[0079] In one of the optional embodiments, the robot includes at least two cameras, and the field of view of at least two cameras constitutes the field of view of the robot; face recognition is performed on the environmental image, including: parallel face recognition is performed on the environmental images captured by each camera, and environmental images with human faces are determined; the robot is controlled to rotate based on the position of the human face so that the human face is within the narrow beam receiving range of the robot to wake up the robot, including: based on the position of the human face, the robot is controlled to rotate in sequence so that the human face is within the narrow beam receiving range of the robot to wake up the robot, and the awakened person is determined based on the acquired audio signal.

[0080] Wherein, when the environment picture includes at least two pictures, face recognition can be performed on each environment picture in parallel to improve processing efficiency. And the face recognized by each environment picture is obtained. Wherein there is no face in the environment picture, there is one face in the environment picture, and there are multiple faces in the environment picture, no subsequent processing is performed on the environment picture without a face, and the position of the face is determined for the environment picture with one face or multiple faces, and the rotation of the robot is subsequently controlled based on the position of the face so that the face of the person speaking is located within the narrow beam receiving range of the robot.

[0081] In one of the optional embodiments, after controlling the robot to rotate based on the position of the face, the steps include obtaining an audio signal received by the robot; converting the audio signal into text information; calling a large model to process the text information to obtain response information corresponding to the audio signal; and outputting the response information.

[0082] After the robot is awakened, it can obtain the audio signal of a person within its narrow beam receiving range and perform subsequent processing, including converting the audio signal into text information, calling a large model to process the text information to obtain response information corresponding to the audio signal, and the response information includes voice or action, etc., which is not specifically limited here.

[0083] Among them combined Figure 4 As shown, the step of calling the large model to process the text information to obtain the response information corresponding to the audio signal may include:

[0084] S402: Obtain scene description information corresponding to the current scene, where the scene description information is obtained by performing attention mechanism processing on the target scene image.

[0085] The scene description information is obtained by performing attention mechanism processing based on the target scene image, and the scene description information is used to describe the attention processing result between any two tokens in the scene. The target scene image is an image of the scene in which the robot is located, and the robot performs image acquisition in real time during operation. In some optional embodiments, the robot may include multiple cameras, and each camera can capture a panoramic image of the scene in which the robot is located. In other embodiments, the robot may include only one camera or a binocular camera, and the field of view of the camera is the field of view of the robot, so that the image captured by the camera can represent the image of the scene in which the robot is located.

[0086] The attention mechanism is the attention mechanism of the decoder in the pre-trained model VLM, which is used to calculate the attention processing results between any two tokens output by the feature extraction network.

[0087] S404: receiving inquiry information.

[0088] The inquiry information is information input by the user, for example, the inquiry information may be “what is included in the scene”, “what are the details of an object in the scene”, etc. The user may make inquiries based on needs.

[0089] After receiving the query information, the query information is processed to obtain a text sequence, for example, the text is segmented to obtain a vector corresponding to each segmented word, and then the vectors are connected to obtain a text sequence corresponding to the query information. In other embodiments, other methods can also be used for processing.

[0090] In addition, it should be noted that if the inquiry information is in the form of voice, in this application, the inquiry information is first converted from voice to text to obtain text information, and then the text information is processed.

[0091] S406: Input the query information and the cached scene description information into the pre-trained model to obtain a response result corresponding to the query information.

[0092] The pre-trained model VLM includes an attention mechanism, wherein the attention mechanism needs to calculate the attention results between any two tokens. In this application, it is assumed that the scene description information includes the attention results between three tokens, namely A, B and C. The scene description information includes the attention results between AB, the attention results between AC and the attention results between BC. Now there is query information. If the corresponding image-text conversion results are to be obtained, it is also necessary to calculate the attention results between the query information Q and the above three tokens, that is, the attention results between AQ, the attention results between BQ and the attention results between CQ.

[0093] In the traditional technology, each time a query is received, the scene image and the query need to be input into the VLM, and the features of the scene image and the text of the query need to be extracted first. Then, the attention mechanism is calculated through transform. Taking the above text as an example, it is necessary to calculate the attention results between AB, the attention results between AC, the attention results between AQ, the attention results between BC, the attention results between BQ, and the attention results between CQ. These attention results need to be calculated each time a query is input, resulting in a large amount of repeated calculations.

[0094] In order to solve the above technical problems, the attention results between AB, AC and BC are stored as scene description information in this application. In this way, in subsequent processing, only the attention results between each token of the scene image and the query information Q need to be calculated. In this application, each token of the scene image also needs to be cached. This can reduce the repeated calculation of the scene description information involved in the query processing summary and improve the processing efficiency.

[0095] The above-mentioned robot question and answer response generation method first obtains the scene image, and obtains the scene description information based on the scene image, and caches it in advance. When making inquiries later, it is directly processed based on the query information and the scene description information without having to calculate the attention result between the scene description information again. This makes full use of the scene description information, reduces the data processing amount of the model, and improves processing efficiency.

[0096] In one of the optional embodiments, in combination Figure 5 As shown, Figure 5The present invention is a flowchart of a scene description information acquisition step in an embodiment. The scene description information acquisition step, i.e., acquiring scene description information corresponding to the current scene, includes: acquiring a target scene image; inputting the target scene image into a pre-trained model, and using the processing result of the attention mechanism in the pre-trained model as the scene description information; and caching the scene description information.

[0097] The pre-trained model VLM is a model that includes robot vision and a large prediction model, wherein the structure of the large prediction model can be a transform structure, so that the attention mechanism of the transform can be utilized, such as the cross-attention mechanism to calculate the cross-attention results of each element in the scene.

[0098] The pre-trained model VLM includes a feature extraction network and a decoder corresponding to a transform. The feature extraction network is used to extract the scene sequence corresponding to the scene image, and then the scene sequence is input into the transform for decoding. The decoding process of the transform includes the processing of the cross-attention mechanism, which can calculate the attention results between different elements in the scene sequence and cache the attention results as scene description information.

[0099] In the above embodiment, scene description information is generated in each new scene, so that when the query is subsequently processed, there is no need to process the scene image, but the scene description information and the query are directly processed, which reduces the repeated calculation of the scene description information and improves the calculation efficiency.

[0100] In one of the optional embodiments, acquiring the target scene image includes: determining the difference between the current scene image and the historical scene image; and when the difference is greater than a threshold, using the current scene image as the target scene image.

[0101] The current scene image is a scene image captured by the robot's camera in real time, and the historical scene image is a scene image previously captured by the robot's camera. The historical scene image may be a target scene image of a previous scene, or any scene image captured from the previous scene to the current time. Optionally, the present application uses the historical scene image as an example of a target scene image of a previous scene for illustration.

[0102] Calculating the difference between the current scene image and the historical scene image may be based on slam. Specifically, calculating the difference between the current scene image and the historical scene image may be calculating the difference in feature points between the current scene image and the historical scene image. For example, the same portion in the current scene image and the historical scene image may be calculated, and then the ratio of the number of pixels in the same portion to the number of pixels in the current scene image may be used as the difference. In other embodiments, the difference may also be determined based on optical flow information, fixed objects, etc., which are not specifically limited here.

[0103] The threshold is preset, and the threshold may be obtained based on experience, such as 75%, etc. In other embodiments, other values ​​may also be selected, and no specific limitation is made here.

[0104] When the difference is greater than the threshold, it is determined that the scene has changed, and the target scene image corresponding to the scene needs to be obtained. Optionally, in order to facilitate processing in this application, the current scene image with a difference greater than the threshold is directly used as the target scene image, that is, the first frame scene image in the new scene.

[0105] In the above embodiment, the scene change can be judged based on the context scene image, thereby ensuring the real-time performance of the target scene image and further ensuring the accuracy of the scene description information.

[0106] In one of the optional embodiments, the target scene image is input into a pre-trained model, and the processing result of the attention mechanism in the pre-trained model is used as the scene description information, including: extracting the scene sequence corresponding to the scene image through the feature extraction network of the pre-trained model; using the scene sequence as the input of the decoder, and using the attention mechanism of the decoder to calculate the attention results between different elements in the scene sequence; and using the attention results as the scene description information.

[0107] Among them, the feature extraction network can be a simple neural network, which can extract the scene sequence corresponding to the scene image. Optionally, the feature extraction network can be connected to a fully connected layer. After the scene sequence is extracted by the scene extraction network, it is then classified through the fully connected layer to complete the prediction of each token, and then the obtained token is input into the decoder as a scene sequence.

[0108] The decoder is transform, and the decoder may include an attention mechanism. The attention mechanism may include a cross-attention mechanism, and the attention result between any two tokens may be calculated through the cross-attention mechanism.

[0109] In the above embodiment, after entering a new scene, the scene description information is first calculated as the basis for subsequent inquiries. In this way, there is no need to calculate the attention results between any two tokens in the scene in subsequent inquiries, which can improve processing efficiency.

[0110] In one of the optional embodiments, the scene description information is cached, including: storing the scene description information in a k-value cache, wherein the index k of the scene description information in the k-value cache is a token, and the value value is an attention result between any two tokens.

[0111] The scene description information includes the attention results between different tokens, where the index can be a token and the value is the attention result between any two tokens.

[0112] In one of the optional embodiments, in combination Figure 3 As shown, Figure 3 The present invention is a processing flow chart for a case where a response result is that an answer corresponding to the query information is not obtained in an embodiment, and the query information and cached scene description information are input into a pre-trained model to obtain a response result corresponding to the query information, including: when the response result is that an answer corresponding to the query information is not obtained, keyword extraction is performed on the query information; image segmentation is performed on the target scene image based on the extracted keywords to obtain a region of interest; and the query information and the region of interest are input into the pre-trained model for processing to obtain a target response result.

[0113] Among them, the response result may include the answer corresponding to the inquiry information, and the answer that cannot be obtained corresponding to the inquiry information. If the response result is the answer corresponding to the inquiry information, the response result can be directly output. If the response result is that the answer corresponding to the inquiry information cannot be obtained, one method can output the answer that cannot be obtained. Another method is to generate response information in combination with context information when the response result is that the answer corresponding to the inquiry information cannot be obtained.

[0114] The keyword extraction of the query information can be performed through a large prediction model. After the target scene image is acquired, the target scene image is processed through the VLM, and the scene elements, that is, the element information corresponding to the token, such as table, chair, etc., can be output.

[0115] By using the big prediction model to extract keywords from query information, keywords can be extracted from the query based on the scene elements output by the VLM, such as the keyword "table".

[0116] The target scene image is segmented based on the extracted keywords, which may be inputting the keywords and the target scene image into a segmentation model, which may be a SAM model or the like, which is not specifically limited here, and the target scene image is segmented based on the keywords by the segmentation model to determine the region of interest corresponding to the query information, and then the region of interest is segmented from the target scene image. This can reduce the amount of image processing.

[0117] Finally, the region of interest and query information are input into the VLM for processing, where the feature extraction network of the VLM is used to extract the image sequence of the region of interest, and the text processing module is used to extract the text features of the query information to obtain the text sequence. The image sequence and text sequence are then input into the VLM decoder, and the attention mechanism of the VLM decoder is used to perform attention processing on the image sequence and text sequence, and finally the target response result is obtained.

[0118] In the above embodiment, when no answer corresponding to the query information is obtained, the contextual query information can be used to segment the scene image to obtain the region of interest, and then VLM processing is performed based on the query information and the scene image to obtain the target response result, thereby making full use of the contextual information to make the robot's response result more accurate.

[0119] It should be understood that, although the steps in the flowcharts involved in the above embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0120] Based on the same inventive concept, the embodiment of the present application also provides a robot awakening device for implementing the robot awakening method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more robot awakening device embodiments provided below can refer to the limitations of the robot awakening method above, and will not be repeated here.

[0121] In an exemplary embodiment, Figure 6As shown, a robot awakening device is provided, including: an environment image acquisition module 601, a face recognition module 602, a face position determination module 603 and a control module 604, wherein:

[0122] The environment image acquisition module 601 is used to acquire the environment image collected by the robot;

[0123] A face recognition module 602 is used to perform face recognition on the environment image;

[0124] A face position determination module 603, configured to determine the position of a face when a face is recognized in at least one of the environment images;

[0125] The control module 604 is used to control the rotation of the robot based on the position of the human face so that the human face is located within the narrow beam receiving range of the robot to wake up the robot.

[0126] In one of the optional embodiments, the face position determination module 603 is specifically configured to determine the position of the face in parallel when at least one face is identified.

[0127] The above-mentioned device also includes: a sorting module, which is used to sort the determined positions of the human faces based on the relative distance between the narrow beam sound receiving range of the robot and the position of the human face.

[0128] In one of the optional embodiments, the above-mentioned device also includes: an open and closed mouth detection module, which is used to segment a face image from an environmental image based on the position of the face; perform open and closed mouth detection on the face image; when the face image has an open or closed mouth, continue to execute the step of controlling the rotation of the robot based on the position of the face; when the face image does not have an open or closed mouth, ignore the position of the face.

[0129] In one of the optional embodiments, the above-mentioned device also includes: a tracking module, which is used to obtain in real time the position of the wake-up person whose face is located within the narrow beam receiving range of the robot; when the position of the wake-up person changes, obtain the first relative posture change of the wake-up person; determine the second relative posture change of the robot based on the first relative posture change; and control the robot to track the wake-up person based on the second relative posture change of the robot.

[0130] In one of the optional embodiments, the robot includes at least two cameras, and the field of view of at least two cameras constitutes the field of view of the robot; the above-mentioned face recognition module 602 is specifically used to perform parallel face recognition on the environmental images collected by each camera, and determine the environmental images in which there are faces.

[0131] The control module 604 is specifically used to control the rotation of the robot in sequence based on the position of the human face, so that the human face is located within the narrow beam receiving range of the robot to wake up the robot, and determine the wake-up person based on the acquired audio signal.

[0132] In one of the optional embodiments, the above-mentioned device also includes: a response module, used to obtain the audio signal received by the robot; convert the audio signal into text information; call the large model to process the text information to obtain response information corresponding to the audio signal; and output the response information.

[0133] Each module in the robot awakening device can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.

[0134] In an exemplary embodiment, a computer device is provided, which may be a robot terminal, and its internal structure diagram may be as shown in FIG. Figure 7 As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC) or other technologies. When the computer program is executed by the processor, a robot wake-up method is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.

[0135] Those skilled in the art will understand that Figure 7The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0136] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the following steps when executing the computer program: acquiring environmental images collected by a robot; performing face recognition on the environmental images; determining the position of the face when a face is recognized in at least one of the environmental images; and controlling the rotation of the robot based on the position of the face so that the face is within the narrow beam receiving range of the robot to wake up the robot.

[0137] In one embodiment, the determination of the position of a human face implemented when the processor executes a computer program includes: in the case where at least one human face is identified, the position of the human face is determined in parallel; when the processor executes the computer program, the following steps are also implemented: based on the relative distance between the robot's narrow beam receiving range and the position of the human face, the determined positions of the human face are sorted.

[0138] In one embodiment, after determining the position of a human face when the processor executes a computer program, the process also includes: segmenting a face image from an environment image based on the position of the face; performing mouth opening and closing detection on the face image; if the face image has an open or closed mouth, continuing to execute the step of controlling the rotation of the robot based on the position of the face; if the face image does not have an open or closed mouth, ignoring the position of the face.

[0139] In one embodiment, after the robot rotation is controlled based on the position of the human face when the processor executes the computer program, it includes: obtaining the position of the wake-up person whose face is located within the narrow beam receiving range of the robot in real time; when the position of the wake-up person changes, obtaining the first relative posture change of the wake-up person; determining the second relative posture change of the robot based on the first relative posture change; and controlling the robot to track the wake-up person based on the second relative posture change of the robot.

[0140] In one embodiment, the cameras on the robot involved when the processor executes the computer program include at least two, and the field of view of the at least two cameras constitutes the field of view of the robot; the face recognition of the environmental image implemented by the processor when executing the computer program includes: parallel face recognition of the environmental images collected by each camera, and determining the environmental image with the human face; the robot rotation control based on the position of the face implemented when the processor executes the computer program so that the human face is within the narrow beam receiving range of the robot to wake up the robot includes: controlling the robot rotation in sequence based on the position of the face so that the human face is within the narrow beam receiving range of the robot to wake up the robot, and determining the wake-up person based on the acquired audio signal.

[0141] In one embodiment, after the processor executes the computer program, the face position-based control of the robot rotation includes: obtaining the audio signal received by the robot; converting the audio signal into text information; calling the large model to process the text information to obtain response information corresponding to the audio signal; and outputting the response information.

[0142] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: obtaining environmental images collected by a robot; performing face recognition on the environmental images; when a face is recognized in at least one of the environmental images, determining the position of the face; and controlling the rotation of the robot based on the position of the face so that the face is within the narrow beam receiving range of the robot to wake up the robot.

[0143] In one embodiment, the method of determining the position of a human face when the computer program is executed by the processor includes: determining the position of the human face in parallel when at least one human face is identified; the computer program also implements the following steps when the processor executes: sorting the determined positions of the human faces based on the relative distance between the robot's narrow beam receiving range and the position of the human face.

[0144] In one embodiment, after the computer program is executed by the processor to determine the position of the face, it also includes: segmenting the face image from the environment image based on the position of the face; performing open and closed mouth detection on the face image; if the face image has an open or closed mouth, continuing to execute the step of controlling the rotation of the robot based on the position of the face; if the face image does not have an open or closed mouth, ignoring the position of the face.

[0145] In one embodiment, after the computer program is executed by the processor, the robot rotation is controlled based on the position of the human face, including: obtaining the position of the wake-up person whose face is located within the narrow beam receiving range of the robot in real time; when the position of the wake-up person changes, obtaining the first relative posture change of the wake-up person; determining the second relative posture change of the robot based on the first relative posture change; and controlling the robot to track the wake-up person based on the second relative posture change of the robot.

[0146] In one embodiment, the cameras on the robot involved when the computer program is executed by the processor include at least two, and the field of view of the at least two cameras constitutes the field of view of the robot; the face recognition of environmental images implemented when the computer program is executed by the processor includes: parallel face recognition of environmental images captured by each camera, and determining environmental images in which human faces exist; the robot rotation control based on the position of the face so that the face is within the narrow beam receiving range of the robot to wake up the robot when the computer program is executed by the processor includes: controlling the robot rotation in sequence based on the position of the face so that the face is within the narrow beam receiving range of the robot to wake up the robot, and determining the wake-up person based on the acquired audio signal.

[0147] In one embodiment, after the computer program is executed by the processor, the robot rotation is controlled based on the position of the human face, including: obtaining an audio signal received by the robot; converting the audio signal into text information; calling a large model to process the text information to obtain response information corresponding to the audio signal; and outputting the response information.

[0148] In one embodiment, a computer program product is provided, including a computer program, which, when executed by a processor, implements the following steps: acquiring environmental images collected by a robot; performing face recognition on the environmental images; determining the position of the face when a face is recognized in at least one of the environmental images; and controlling the rotation of the robot based on the position of the face so that the face is within the narrow beam receiving range of the robot to wake up the robot.

[0149] In one embodiment, the method of determining the position of a human face when the computer program is executed by the processor includes: determining the position of the human face in parallel when at least one human face is identified; the computer program also implements the following steps when the processor executes: sorting the determined positions of the human faces based on the relative distance between the robot's narrow beam receiving range and the position of the human face.

[0150] In one embodiment, after the computer program is executed by the processor to determine the position of the face, it also includes: segmenting the face image from the environment image based on the position of the face; performing open and closed mouth detection on the face image; if the face image has an open or closed mouth, continuing to execute the step of controlling the rotation of the robot based on the position of the face; if the face image does not have an open or closed mouth, ignoring the position of the face.

[0151] In one embodiment, after the computer program is executed by the processor, the robot rotation is controlled based on the position of the human face, including: obtaining the position of the wake-up person whose face is located within the narrow beam receiving range of the robot in real time; when the position of the wake-up person changes, obtaining the first relative posture change of the wake-up person; determining the second relative posture change of the robot based on the first relative posture change; and controlling the robot to track the wake-up person based on the second relative posture change of the robot.

[0152] In one embodiment, the cameras on the robot involved when the computer program is executed by the processor include at least two, and the field of view of the at least two cameras constitutes the field of view of the robot; the face recognition of environmental images implemented when the computer program is executed by the processor includes: parallel face recognition of environmental images captured by each camera, and determining environmental images in which human faces exist; the robot rotation control based on the position of the face so that the face is within the narrow beam receiving range of the robot to wake up the robot when the computer program is executed by the processor includes: controlling the robot rotation in sequence based on the position of the face so that the face is within the narrow beam receiving range of the robot to wake up the robot, and determining the wake-up person based on the acquired audio signal.

[0153] In one embodiment, after the computer program is executed by the processor, the robot rotation is controlled based on the position of the human face, including: obtaining an audio signal received by the robot; converting the audio signal into text information; calling a large model to process the text information to obtain response information corresponding to the audio signal; and outputting the response information.

[0154] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0155] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.

[0156] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0157] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A robot awakening method, characterized in that: The method comprises: Get the environment pictures collected by the robot; Performing face recognition on the environment picture; When a face is recognized in at least one of the environment images, determining a position of the face; The robot is controlled to rotate based on the position of the human face so that the human face is located within the narrow beam sound receiving range of the robot to wake up the robot.

2. The method according to claim 1, characterized in that The determining the position of the face comprises: When at least one face is identified, the position of the face is determined in parallel; Before controlling the robot to rotate based on the position of the human face, the method further includes: The determined positions of the human faces are sorted based on the relative distance between the narrow beam sound receiving range of the robot and the position of the human face.

3. The method according to claim 2, characterized in that After determining the position of the face, the method further includes: Segmenting the face image from the environment image based on the position of the face; Performing open and closed mouth detection on the face image; In the case where the face image has an open or closed mouth, continuing to execute the step of controlling the robot to rotate based on the position of the face; When the face image does not have an open or closed mouth, the position of the face is ignored.

4. The method according to any one of claims 1 to 3, characterized in that: After controlling the robot to rotate based on the position of the human face, the method further comprises: Real-time acquisition of the position of the wake-up person whose face is within the narrow beam receiving range of the robot; When the position of the awakening person changes, acquiring a first relative posture change of the awakening person; Determining a second relative posture change of the robot based on the first relative posture change; Based on the change of the second relative posture of the robot, the robot is controlled to track and wake up the person.

5. The method according to any one of claims 1 to 3, characterized in that: The robot includes at least two cameras, and the field of view of at least two cameras constitutes the field of view of the robot; performing face recognition on the environment image includes: Performing parallel face recognition on the environmental images collected by the cameras, and determining environmental images containing human faces; The controlling the robot to rotate based on the position of the human face so that the human face is located within the narrow beam sound receiving range of the robot to wake up the robot includes: The robot is controlled to rotate in sequence based on the position of the human face so that the human face is located within the narrow beam sound receiving range of the robot to wake up the robot, and the awakened person is determined based on the acquired audio signal.

6. The method according to any one of claims 1 to 3, characterized in that: After controlling the robot to rotate based on the position of the human face, the method further comprises: Obtain the audio signal received by the robot; Converting the audio signal into text information; Calling the large model to process the text information to obtain response information corresponding to the audio signal; The response information is output.

7. A robot awakening device, characterized in that: The device comprises: The environment image acquisition module is used to obtain the environment images collected by the robot; A face recognition module, used for performing face recognition on the environment image; A face position determination module, used to determine the position of the face when a face is recognized in at least one of the environment images; The control module is used to control the rotation of the robot based on the position of the human face so that the human face is located within the narrow beam receiving range of the robot to wake up the robot.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Intelligent robot tracking method and tracking device based on artificial intelligence

    CN105116994A

  • Sound pickup method and device

    CN109688512A

  • Voice interaction wake-up-free method and device

    CN112634895A

  • Man-machine conversation method, electronic equipment and computer readable storage medium

    CN112634911A

  • Control method and device of vehicle-mounted intelligent equipment, vehicle and storage medium

    CN116080565A