Humanoid robot information input method, system and device and storage medium
By obtaining the trigger signal of the humanoid robot, determining the personnel level and performing information entry at the corresponding level, and using multi-modal sensors to obtain identity information, the problem of a single information entry method of humanoid robot is solved, and efficient, accurate and safe information management of multi-level personnel is achieved.
Patent Information
- Application Number
- CN202510326378.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, the information entry method of humanoid robots is single, and it is difficult to distinguish the information entry requirements and permissions of personnel of different levels, resulting in low entry efficiency and security risks, and insufficient accuracy and stability of information entry.
By obtaining the trigger signal based on the humanoid robot, determining the personnel level according to the signal type and signal source, and performing information entry and guidance at the corresponding level, using multi-modal sensors to obtain identity information, and combining multi-sensor data verification, it realizes accurate identification and information entry of multi-level personnel.
It improves the efficiency and security of information entry, improves the convenience and accuracy of the entry process, and enhances the user experience and the intelligent management capabilities of the robot.
Smart Images

Figure CN120375465A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robots, and in particular to a method, system, device and storage medium for inputting information of a humanoid robot. Background Art
[0002] With the wide application of humanoid robots in fields such as home, office, and service, the efficient input and management of personnel information have become crucial. In related technologies, for the input of information by humanoid robots, a single program and device are mostly used for input, and the obtained information is messy, affecting the input efficiency. Summary of the Invention
[0003] The purpose of the present invention is to solve at least to a certain extent one of the technical problems existing in the prior art.
[0004] To this end, the purpose of the present invention is to provide an efficient method, system, device and storage medium for inputting information of a humanoid robot.
[0005] In order to achieve the above technical purpose, on the one hand, an embodiment of the present invention provides a method for inputting information of a humanoid robot, including the following steps: obtaining a trigger signal based on a target humanoid robot; determining a personnel level according to the signal type and signal source of the trigger signal; and according to the personnel level, performing information input guidance at the corresponding level and inputting the information into the target humanoid robot. In this application, the personnel level is determined by the trigger signal, and different input guides are set according to the personnel level for information input. This application can determine the input information according to the personnel level, which is beneficial to improving the input efficiency.
[0006] In an embodiment of the present invention, the signal type includes a voice signal, a gesture signal, and a touch signal. The determining the personnel level according to the signal type and signal source of the trigger signal includes:
[0007] If the voice signal includes a preset level description, determining the personnel level according to the level corresponding to the preset level description;
[0008] Or, processing the gesture signal to determine the gesture complexity corresponding to the gesture signal, and determining the personnel level according to the gesture complexity;
[0009] Or, determining the touch area corresponding to the touch signal, and determining the personnel level according to the touch area.
[0010] In an embodiment of the present invention, the signal type includes a voice signal, a gesture signal, and a touch signal. The determining the personnel level according to the signal type and signal source of the trigger signal includes:
[0011] Set the priorities of the voice signal, the gesture signal, and the touch signal, and determine the personnel level according to the signal type in the order from high to low or from low to high of the priorities;
[0012] Or;
[0013] Set the signal confidence levels of the voice signal, the gesture signal, and the touch signal, and determine the corresponding preset categories respectively according to the voice signal, the gesture signal, and the touch signal;
[0014] Determine the personnel category according to the preset category and the corresponding signal confidence level.
[0015] In an embodiment of the present invention, the signal type includes a voice signal, a gesture signal, and a touch signal. The determining of the personnel level according to the signal type and the signal source of the trigger signal includes:
[0016] Determine several preset levels according to the voice signal, the gesture signal, and the touch signal;
[0017] If the preset levels corresponding to at least two of the voice signal, the gesture signal, and the touch signal are the first level, determine the personnel level as the first level;
[0018] Or;
[0019] If the preset levels corresponding to at least two of the voice signal, the gesture signal, and the touch signal are different, re-collect the trigger signal through information reminder.
[0020] In an embodiment of the present invention, the signal source includes a video acquisition device and a voice acquisition device. The determining of the personnel level according to the signal type and the signal source of the trigger signal includes:
[0021] Determine the voice positioning according to the time difference of the voice signal reaching each voice acquisition device;
[0022] Determine the video positioning according to the direction and visual information of the video acquisition device;
[0023] Determine the position information according to the voice positioning and the video positioning;
[0024] Determine the personnel level according to the position information and the position association rule.
[0025] In an embodiment of the present invention, the method further includes:
[0026] Match the gesture signal with the signals in the standard gesture library through a matching algorithm, and determine the personnel level according to the matching result;
[0027] Alternatively, extract the gesture features from the gesture signal, and process the gesture features through a machine learning model to determine the gesture complexity and the personnel level.
[0028] In an embodiment of the present invention, the method further includes:
[0029] Perform information entry operations at the corresponding level through mutual complementation and verification of multiple sensors to obtain the entered information.
[0030] On the other hand, an embodiment of the present invention provides a humanoid robot information entry system, including:
[0031] A first module for obtaining a trigger signal based on a target humanoid robot;
[0032] A second module for determining the personnel level according to the signal type and signal source of the trigger signal;
[0033] A third module for performing information entry guidance at the corresponding level according to the personnel level and entering the information into the target humanoid robot.
[0034] On the other hand, an embodiment of the present invention provides a humanoid robot information entry device, including:
[0035] At least one processor;
[0036] At least one memory for storing at least one program;
[0037] When the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned humanoid robot information entry method.
[0038] On the other hand, an embodiment of the present invention provides a storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to implement the above-mentioned humanoid robot information entry method when executed by the processor.
[0039] The embodiments of the present application at least include the following beneficial effects: The method provided by the embodiments of the present invention includes: obtaining a trigger signal based on a target humanoid robot; determining the personnel level according to the signal type and signal source of the trigger signal; performing information entry guidance at the corresponding level according to the personnel level and entering the information into the target humanoid robot. The present application determines the personnel level through the trigger signal, sets different entry guides according to the personnel level, and performs information entry. The present application can determine the entered information according to the personnel level, which is beneficial to improving the entry efficiency. Description of the Drawings
[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the accompanying drawings of the relevant technical solutions in the embodiments of the present invention or the prior art. It should be understood that the accompanying drawings in the following introduction are only for clearly expressing some embodiments of the technical solutions in the present invention. For those skilled in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0041] Figure 1 A flowchart of an embodiment of the humanoid robot information entry method provided by the present invention;
[0042] Figure 2 A flowchart of an embodiment of the humanoid robot face information entry method provided by the present invention;
[0043] Figure 3 A flowchart of another embodiment of the humanoid robot information entry method provided by the present invention;
[0044] Figure 4 A structural diagram of an embodiment of the humanoid robot information entry system provided by the present invention;
[0045] Figure 5 A structural diagram of another embodiment of the humanoid robot information entry system provided by the present invention;
[0046] Figure 6 A structural diagram of an embodiment of the humanoid robot information entry device provided by the present invention. Detailed Embodiments
[0047] The following details the embodiments of the present invention. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention. For the step numbers in the following embodiments, they are only set for the convenience of explanation and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0048] With the wide application of humanoid robots in fields such as home, office, and service, the efficient input and management of information for different levels of personnel (such as users, administrators, etc.) have become crucial. However, there are many deficiencies in the existing humanoid robots in terms of personnel information input: First, the input method is single, mostly relying on a single interaction mode, such as only operating through a touch screen or simple voice commands, unable to make full use of the multi-modal interaction capabilities of humanoid robots, making the input process less flexible and convenient. Second, it is difficult to distinguish the information input requirements and permissions of different levels of personnel, lacking pertinence in the input process and permission management, which may lead to information security risks or low input efficiency. For example, an administrator needs to input system configuration information with higher permissions, while an ordinary user only needs to input basic usage preference information, but the existing technology cannot efficiently distinguish and process these different requirements. Third, the accuracy and stability of information input are insufficient, easily affected by environmental factors (such as changes in light affecting face recognition and noisy environments interfering with voice recognition), affecting the user experience and the normal operation of the robot. Therefore, there is an urgent need for a multi-level personnel information input method that can solve the above problems.
[0049] The following will describe in detail the humanoid robot information input method and system according to the embodiments of the present invention with reference to the accompanying drawings. First, the humanoid robot information input method according to the embodiments of the present invention will be described with reference to the accompanying drawings.
[0050] Refer to Figure 1 , in the embodiments of the present invention, a humanoid robot information input method is provided. The humanoid robot information input method in the embodiments of the present invention can be applied to a terminal, or to a server, or can also be software running on a terminal or a server, etc. The terminal can be a tablet computer, a notebook computer, a desktop computer, etc., but is not limited thereto. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The humanoid robot information input method in the embodiments of the present invention mainly includes the following steps:
[0051] S100: Obtain a trigger signal based on the target humanoid robot;
[0052] S200: Determine the personnel level according to the signal type and signal source of the trigger signal;
[0053] S300: According to the personnel level, execute the information input guidance for the corresponding level and input the information into the target humanoid robot.
[0054] In some possible embodiments, the signal type is the manifestation form of the signal, and the signal source is the acquisition method of the signal. The target humanoid robot is the humanoid robot for which information is to be entered. According to the personnel level, the present application determines the corresponding entry guidance and executes information entry in different orders, can determine the order of information entry according to different personnel levels, is beneficial to improving the information security performance, and improves the entry efficiency. The present application discloses a method for entering multi-level personnel information of a humanoid robot, an electronic device, and a storage medium, relating to the technical field of humanoid robots. The method includes: obtaining a trigger signal, determining the target entry personnel level according to the signal type and source, such as user level or administrator level; obtaining the identity information of the corresponding personnel through a multi-modal sensor, such as obtaining a face image based on a vision sensor and collecting voice features using a voice sensor; and completing the entry of personnel information based on the obtained information. This method simplifies the process of entering multi-level personnel information, combines multi-modal perception to achieve accurate identification and information entry of different-level personnel, improves the interaction convenience and intelligence of the humanoid robot in personnel information management, and effectively enhances the user experience.
[0055] In an embodiment of the present invention, the signal type includes a voice signal, a gesture signal, and a touch signal. Determining the personnel level according to the signal type and source of the trigger signal includes:
[0056] If the voice signal includes a preset level description, determine the personnel level according to the level corresponding to the preset level description;
[0057] Or, process the gesture signal, determine the gesture complexity corresponding to the gesture signal, and determine the personnel level according to the gesture complexity;
[0058] Or, determine the touch area corresponding to the touch signal, and determine the personnel level according to the touch area.
[0059] In the present application, the preset level description may be words related to the personnel level such as "administrator" and "user" included in the voice signal. The present application determines the gesture complexity through the gesture signal, and then determines the personnel level according to the relationship between the gesture complexity and the threshold. The touch area in the present application may be a physical touch area or a virtual permission area.
[0060] In an embodiment of the present invention, the signal type includes a voice signal, a gesture signal, and a touch signal. Determining the personnel level according to the signal type and source of the trigger signal includes:
[0061] Set the priorities of the voice signal, the gesture signal, and the touch signal, and determine the personnel level according to the signal type in the order from high to low or from low to high of the priorities;
[0062] Or;
[0063] Set the signal confidence levels of voice signals, gesture signals, and touch signals, and respectively determine the corresponding preset categories according to the voice signals, gesture signals, and touch signals;
[0064] Determine the personnel category according to the preset category and the corresponding signal confidence level.
[0065] For the voice signals, gesture signals, and touch signals in this application, each signal can determine a personnel level. The signal types can be prioritized, and the personnel level corresponding to the signal type with a higher priority is used as the final personnel level.
[0066] In an embodiment of the present invention, the signal types include voice signals, gesture signals, and touch signals. According to the signal type and signal source of the trigger signal, determining the personnel level includes:
[0067] Determine a number of preset levels according to the voice signals, gesture signals, and touch signals;
[0068] If the preset levels corresponding to at least two of the voice signal, gesture signal, and touch signal are the first level, determine the personnel level as the first level;
[0069] Or;
[0070] If the preset levels corresponding to at least two of the voice signal, gesture signal, and touch signal are different, re-collect the trigger signal through information reminder.
[0071] In an embodiment of the present invention, the signal sources include video acquisition devices and voice acquisition devices. According to the signal type and signal source of the trigger signal, determining the personnel level includes:
[0072] Determine the voice positioning according to the time difference of the voice signal arriving at each voice acquisition device;
[0073] Determine the video positioning according to the direction and visual information of the video acquisition device;
[0074] Determine the location information according to the voice positioning and video positioning;
[0075] Determine the personnel level according to the location information and the location association rule.
[0076] In an embodiment of the present invention, the method further includes:
[0077] Match the gesture signal with the signals in the standard gesture library through a matching algorithm, and determine the personnel level according to the matching result;
[0078] Alternatively, extract the gesture features in the gesture signal, and process the gesture features through a machine learning model to determine the gesture complexity and the personnel level.
[0079] In one embodiment of the present invention, the method further includes:
[0080] Performing information entry operations at corresponding levels through mutual complementation and verification of multiple sensors to obtain the entered information.
[0081] The following uses a specific embodiment to introduce the information entry method proposed in this application in detail:
[0082] This application proposes a multi-level personnel information entry method: when a trigger signal is obtained, based on the type of the trigger signal (in one embodiment, through voice commands, gesture actions, touch operations) and the signal source (in one embodiment, the direction of the voice signal source received by the built-in microphone of the humanoid robot, the position where the gesture action is initiated captured by the camera), determine the target entry personnel level, that is, the user level or the administrator level. Based on the multi-modal sensors carried by the humanoid robot, obtain the identity information of the corresponding level personnel. In one embodiment, as shown in Figure 2 , use a vision sensor to obtain a face image, and determine the facial features of the person by processing and analyzing the image; collect voice features through a voice sensor, and analyze information such as the timbre, intonation, and speech rate of the voice. At the same time, other sensors can also be combined. In one embodiment, fingerprint information is obtained through a fingerprint sensor to further improve the accuracy of identity recognition. Complete personnel information entry based on the obtained personnel identity information. During the entry process, classify and manage the information of different level personnel to provide data support for subsequent permission control and personalized services.
[0083] This application proposes to judge whether the trigger signal contains clear personnel level information (in one embodiment, keywords such as "administrator entry" or "user entry" are included in the voice command, and the touch operation is performed in a specific permission area, etc.). If so, directly determine the target entry personnel level according to this information; otherwise, determine the target entry personnel level based on the type and source of the trigger signal.
[0084] The specific method for determining the entry level based on multi-modal signals proposed in this application: obtain the sound localization result by calculating the time difference of the voice signal arriving at multiple microphones of the humanoid robot, combine the rotation direction of the robot's head and visual information to determine the position of the person initiating the information entry, and then determine the target entry personnel level according to the preset association rule between the position and the personnel level. In one embodiment, an entry request initiated within a specific management area is likely to be at the administrator level; while a request initiated in a normal usage scenario is defaulted to the user level. For an entry request triggered by a gesture action, judge according to the preset corresponding relationship between the gesture action and the personnel level.
[0085] One of the embodiments sets the trigger action for administrator privileges through a specific combination of complex gestures, while simple and common gestures correspond to ordinary user operations.
[0086] The present application proposes to improve the input accuracy by combining multi-sensor data: when obtaining the personal identity information, multiple sensor data complement and verify each other. In one of the embodiments, when using a vision sensor to obtain a face image, the data of the ambient light sensor is combined to adjust the brightness and contrast of the image, improving the accuracy of face feature extraction; when a voice sensor collects voice features, the data of the noise sensor is referred to, and the voice collection parameters are automatically adjusted in a noisy environment to ensure the reliability of the voice features.
[0087] Information input guidance and quality control: After determining the level of the target input person, the person is guided for information input through various means such as voice, screen display, and light indication. In one of the embodiments, clear input step prompts are displayed on the screen, and the current input progress and precautions are informed through voice. Feature extraction is performed based on the obtained personal identity information, and quality calculation is performed according to the extracted features. When the quality calculation result meets the standard, it is prompted that the personal information input is completed; if not, the user is guided to re-enter the information or adjust the input method.
[0088] In implementing the technical solution of the present application, first, the level of the target input person is determined based on the trigger signal, and then the identity information of the corresponding person is obtained through multi-modal sensors to complete the information input. This method simplifies the information input process for multi-level personnel, realizes intelligent identification and information input for different levels of personnel by combining multi-modal perception, improves the interaction convenience and intelligence level of the humanoid robot in personnel information management, makes the personnel information input interaction more concise and natural, and effectively enhances the user experience. At the same time, through the fusion and verification of multi-sensor data, the accuracy and stability of information input are improved, ensuring the efficient operation and information security of the humanoid robot.
[0089] Specifically:
[0090] Refer to Figure 3 and Figure 4As shown in the figure, the present application provides trigger signal acquisition and personnel level determination: the humanoid robot constantly monitors the surrounding environment and starts the personnel information entry process when it receives a trigger signal. The types of trigger signals are diverse. For example, for voice commands, such as voices with specific keyword combinations like "enter my information" or "administrator enter information", the humanoid robot receives the voice signal through a built-in microphone; for gesture actions, such as specific wave or fist-clenching actions, the visual sensor carried by the robot captures the gestures; for touch operations, such as clicks and swipes in specific touch areas of the humanoid robot, the touch sensor records the relevant operations. If the trigger signal contains clear personnel level information, such as the keyword "administrator entry" in the voice command, the robot directly determines it as an administrator level entry; if not, it is determined according to the signal source. For example, for voice signals, the sound source is located by analyzing the time difference of the sound reaching multiple microphones, and combined with the robot's head orientation and visual information, the position of the person is judged. If an entry request is detected in the office area where administrators often interact, it is most likely determined as the administrator level; if in the common activity area of the home, it tends to be the user level. When a gesture action is triggered, specific complex gestures are preset as administrator permissions, and simple and common gestures correspond to user operations.
[0091] Among them, the specific logic for judging the trigger level of gesture actions. One embodiment: comparison with the annotation library. A standard gesture library is established in advance. In one embodiment, through the administrator gesture library and the user gesture library, through template matching. In one embodiment, the dynamic time warping (DTW) and convolutional neural network are used to compare the gestures collected in real time with the templates in the library to judge the belonging category.
[0092] This solution is applicable to fixed gesture scenarios.
[0093] Another embodiment: complexity scoring.
[0094] Extract gesture features, which include the action trajectory length, speed change, and joint angle complexity in the embodiment. Through an algorithm, in one embodiment, a machine learning model is used to quantify the complexity, and a threshold is set to distinguish simple / complex gestures.
[0095] The solution has strong adaptability and does not require maintaining a template library.
[0096] Exemplarily, for the embodiment using template matching, it is necessary to describe in the document the construction method of the gesture library (such as collecting the 3D coordinate sequence of the specific gestures of the administrator) and the matching algorithm (such as the DTW algorithm). In another embodiment, if complexity scoring is adopted, it is necessary to define specific indicators (such as the number of joint points included in the gesture, the entropy value of the motion trajectory) and the scoring threshold (such as a score ≥ 80 is a complex gesture).
[0097] The present application also distinguishes personnel levels through touch operations.
[0098] One of the embodiments: Physical area division:
[0099] Set a physical touch area on the surface of the robot (for example, a certain area on the back is exclusive to the administrator), and locate the contact position through a touch sensor. Advantage of the solution: Simple hardware implementation and high security.
[0100] One of the embodiments: Screen interaction area division:
[0101] Divide a virtual permission area on the robot display screen (for example, the lower right corner of the screen is the administrator operation area), and judge the level through the touch coordinates. Advantage of the solution: High flexibility and can be dynamically adjusted.
[0102] For the above two embodiments of the touch permission area, secondary verification still needs to be combined. One of the embodiments is a password, and the other embodiment is a fingerprint to ensure the legitimacy of the permission and avoid misoperation.
[0103] When the voice signal, gesture signal, and touch signal in the type information are triggered simultaneously, how to determine the final personnel level? One of the embodiments proposed in this application is through priority sorting:
[0104] Set the signal priority. In one of the embodiments, the voice is higher than the gesture and higher than the touch, and the high-priority signal covers the low-priority signal.
[0105] One of the embodiments: Weight comprehensive calculation:
[0106] According to the signal confidence level, in one of the embodiments, the voice matching degree is 90% and the gesture matching degree is 70% to calculate the final result by weighted calculation.
[0107] One of the embodiments: Context correlation verification:
[0108] Combined with the environmental information, in one of the embodiments, the trigger position and time are used to judge the rationality of the signal. For example, if the voice and gesture are triggered simultaneously in the administrator area, it is preferentially judged as an administrator.
[0109] This application also provides conflict handling logic:
[0110] Example 1: When the voice command is "Administrator entry" and the gesture is a complex action, it is directly judged as the administrator level.
[0111] Example 2: If the voice and gesture signals conflict, such as the voice is "User entry" but the gesture is an administrator gesture, start secondary verification, such as fingerprint recognition.
[0112] This application can also add a prompt mechanism when signals conflict. In one of the embodiments, a voice prompt is "Multiple operation requests are detected. Please confirm the permission", and the user is guided to trigger again.
[0113] On the other hand, this application provides for collecting identity information using multi-modal sensors: After determining the personnel level, multi-modal sensors are used to collect identity information. In terms of visual sensors, the robot's camera captures the facial image of the person. To ensure image quality, the shooting parameters are adjusted in combination with the data of the ambient light sensor. When the light is dim, the exposure time is automatically increased or the fill light is turned on; when the light is too bright, the exposure is reduced. After the image is collected, advanced image recognition algorithms are used to extract facial feature points. In one embodiment, this is done by analyzing the contours and relative positions of the eyes, nose, mouth, etc. The voice sensor collects the person's voice, and the data of the noise sensor is referred to during the collection process. If in a noisy environment, the microphone gain is automatically increased or a noise reduction algorithm is used to remove background noise interference and extract features such as the timbre, intonation, and speech rate of the voice. For scenarios that require more precise identity verification, such as when an administrator enters key system information, the fingerprint sensor is enabled to collect fingerprint information, and identity confirmation is performed by comparing fingerprint feature points.
[0114] This application provides information entry guidance and quality control: During the information collection process, the humanoid robot guides the person's operation in various ways. Clear entry step prompts are displayed on the screen. In one embodiment, voice prompts such as "Please align your face with the camera and keep your head stable" and "Please clearly state your information" are used; at the same time, in combination with the voice prompts, the current entry progress is informed. In one embodiment, it is "Collecting your voice information, please wait a moment". After the information is collected, its quality is calculated. The facial image quality calculation comprehensively considers factors such as image clarity and facial feature integrity. If the clarity is lower than the set threshold or some features are blocked, a prompt to re-collect is given; the voice quality calculation is based on the recognizability and signal strength of the voice. If the voice is unclear or the signal is too weak, guidance to re-record is provided. Only when the quality of each piece of information meets the standards is the information entry completed and a success prompt is given.
[0115] This application provides for classified storage and management of information: The entered personnel information is classified and stored in the robot's memory according to the level. User information is stored in the ordinary user data area, including basic identity information, usage preferences, etc.; administrator information is stored in the management data area with higher permissions, which in addition to basic information also includes system management permissions, configuration information, etc. The memory adopts a secure storage architecture, and strict access permissions are set for the management data area. Only authorized administrator operations can access and modify it to ensure information security.
[0116] The present application provides specific application scenario examples: In a home scenario, a user hopes to enter personal information so that the robot can provide personalized services. The user utters the voice command "Enter my information", and the robot determines it as user-level entry through voice recognition and position judgment. The camera captures the user's facial image, the voice sensor captures voice information, and the screen prompts the user to adjust the position and speaking manner to ensure the information quality. After the capture is completed and the quality meets the standard, the robot stores the user information in the corresponding area, and subsequent personalized greetings, entertainment recommendations, etc. can be provided based on this information. In an office scenario, an administrator needs to enter new system management information. The administrator makes a specific gesture, and after the robot recognizes it, it determines it as administrator-level entry and guides the administrator to perform fingerprint verification, facial recognition, and enter system management-related information. After the verification is passed and the information entry is completed, the administrator can manage and configure the system permissions, task allocation, etc. of the robot to ensure the efficient operation of the robot in the office environment.
[0117] On the other hand, referring to Figure 5 as shown, an embodiment of the present invention provides a humanoid robot information entry system, including:
[0118] A first module 510, configured to obtain a trigger signal based on a target humanoid robot;
[0119] A second module 520, configured to determine the personnel level according to the signal type and signal source of the trigger signal;
[0120] A third module 530, configured to perform information entry guidance at the corresponding level according to the personnel level and enter the information into the target humanoid robot.
[0121] It can be seen that the content in the above method embodiments is applicable to the system embodiments of the present invention. The functions specifically implemented by the system embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0122] Referring to Figure 6 , an embodiment of the present invention provides a humanoid robot information entry device, including:
[0123] At least one processor 610;
[0124] At least one memory 620, configured to store at least one program;
[0125] When at least one program is executed by at least one processor 610, a humanoid robot information entry method implemented by at least one processor 610 is enabled.
[0126] Similarly, the content in the above method embodiments is applicable to the embodiments of this apparatus. The functions specifically implemented by the embodiments of this apparatus are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0127] An embodiment of the present invention further provides a computer-readable storage medium, in which a program executable by a processor is stored. The program executable by the processor is used to execute the above-mentioned humanoid robot information entry method when executed by the processor.
[0128] Similarly, the content in the above method embodiments is applicable to the embodiments of this storage medium. The functions specifically implemented by the embodiments of this storage medium are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0129] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order mentioned in the operation diagrams. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. Additionally, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are foreseeable, in which the order of various operations is changed and the sub-operations described as part of a larger operation are executed independently.
[0130] Furthermore, although the present invention has been described in the context of functional modules, it should be understood that unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It can also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. More precisely, considering the attributes, functions, and internal relationships of the various functional modules in the apparatus disclosed herein, the actual implementation of the module will be understood within the ordinary skills of an engineer. Therefore, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation using ordinary skills. It can also be understood that the specific concepts disclosed are illustrative only and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0131] If the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several programs for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0132] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as an ordered list of executable programs for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by a program execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can retrieve and execute programs from the program execution system, apparatus, or device), or in conjunction with these program execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with a program execution system, apparatus, or device.
[0133] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection part with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or, if necessary, other suitable processing, and then stored in a computer memory.
[0134] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable program execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0135] In the foregoing description of this specification, descriptions with reference to the terms "one embodiment / example", "another embodiment / example" or "certain embodiments / examples", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0136] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.
[0137] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present invention.
Claims
1. A method for inputting information of a humanoid robot, characterized in that, Including the following steps: Obtain a trigger signal based on the target humanoid robot; Determine the personnel level according to the signal type and signal source of the trigger signal; According to the personnel level, perform information entry guidance at the corresponding level and enter the information into the target humanoid robot.
2. The method for inputting information of a humanoid robot according to claim 1, wherein The signal type includes voice signals, gesture signals, and touch signals. Determining the personnel level according to the signal type and signal source of the trigger signal includes: If the voice signal includes a preset level description, determine the personnel level according to the level corresponding to the preset level description; Or, process the gesture signal to determine the gesture complexity corresponding to the gesture signal, and determine the personnel level according to the gesture complexity; Or, determine the touch area corresponding to the touch signal, and determine the personnel level according to the touch area.
3. The method for inputting humanoid robot information according to claim 1, wherein The signal type includes voice signals, gesture signals, and touch signals. Determining the personnel level according to the signal type and signal source of the trigger signal includes: Set the priorities of the voice signal, the gesture signal, and the touch signal, and determine the personnel level according to the signal type in the order from high to low or from low to high of the priorities; Or; Set the signal confidence levels of the voice signal, the gesture signal, and the touch signal, and respectively determine the corresponding preset categories according to the voice signal, the gesture signal, and the touch signal; Determine the personnel category according to the preset category and the corresponding signal confidence level.
4. The method for inputting information of a humanoid robot according to claim 1, wherein The signal type includes voice signals, gesture signals, and touch signals. Determining the personnel level according to the signal type and signal source of the trigger signal includes: Determine a number of preset levels according to the voice signal, the gesture signal, and the touch signal; If the preset levels corresponding to at least two of the voice signal, the gesture signal, and the touch signal are the first level, determine the personnel level as the first level; Or; If the preset levels corresponding to at least two of the voice signal, the gesture signal, and the touch signal are different, re-collect the trigger signal through information reminder.
5. The method for inputting humanoid robot information according to claim 1, wherein The signal source includes a video acquisition device and a voice acquisition device. Determining the personnel level according to the signal type and signal source of the trigger signal includes: Determine the voice positioning according to the time difference of the voice signal reaching each voice acquisition device; Determine the video positioning according to the direction and visual information of the video acquisition device; Determine the position information according to the voice positioning and the video positioning; Determine the personnel level according to the position information and the position association rule.
6. The method for inputting information of a humanoid robot according to claim 2, wherein The method further includes: Match the gesture signal with the signals in the standard gesture library through a matching algorithm, and determine the personnel level according to the matching result; Or, extract the gesture features in the gesture signal, and process the gesture features through a machine learning model to determine the gesture complexity and the personnel level.
7. The method for inputting information of a humanoid robot according to claim 1, wherein, The method further includes: Perform information entry operations at the corresponding level through multi-sensor mutual complementary verification to obtain the entered information.
8. A humanoid robot information entry system, characterized in that, Including: The first module is used to obtain a trigger signal based on the target humanoid robot; A second module, configured to determine a personnel level according to the signal type and signal source of the trigger signal; A third module, configured to execute an information entry guidance corresponding to the personnel level and enter information into the target humanoid robot according to the personnel level.
9. An information input device for a humanoid robot, characterized in that, Comprising: At least one processor; At least one memory, configured to store at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the humanoid robot information entry method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that, The program executable by the processor is used to implement the humanoid robot information entry method according to any one of claims 1 to 7 when executed by the processor.
Citation Information
Patent Citations
Device control method, apparatus and computer readable storage medium
CN107678287A
Intelligent reception service method and intelligent reception service system
CN108038421A
Face recognition method and equipment and computer readable storage medium
CN111079791A
Input system for intelligent robot
CN113518065A
Cited By
Training method of humanoid robot language recognition model and language recognition method
CN121260149A