A scenario-based robot interaction control method and system
By analyzing the busyness level of scenes and adjusting the streamlining of voice interaction, the problem of too much irrelevant content in robot voice interaction is solved, and the interaction efficiency and user experience are improved.
Patent Information
- Application Number
- CN202510423174.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The voice content output by existing robots in voice interaction often contains a large number of elements that are not directly related to the interactive content, such as tone words and greeting words, which leads to inefficient interaction, especially in busy scenarios.
By scanning to obtain the panoramic image of the application scene and shooting to obtain the wide-angle image of the interaction direction area, extract personnel feature information, analyze the busy level of the scene, and adjust the succinctness of the voice interaction according to the busy level to reduce the output of irrelevant content.
In busy scenarios, the robot can output more concise and direct voice content, improve the speed of users to obtain key information, and significantly improve the efficiency and user experience of voice interaction.
Smart Images

Figure CN119927928B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voice interaction, and more particularly, to a scene-based robot interaction control method and system. Background Art
[0002] With the rapid development of artificial intelligence technology, robots are increasingly widely used in various fields. As one of the important ways of human-computer interaction, voice interaction has greatly facilitated the communication between people and robots. However, current robots still have significant deficiencies in voice interaction.
[0003] In the existing voice interaction process, the voice output by the robot is often fixedly added with a large number of elements that are not directly related to the interaction content, such as filler words "um", "ah", and greetings "Hello", "How's your day". Although these additional contents can, to a certain extent, create a more natural and friendly interaction atmosphere, they seriously affect the efficiency of voice interaction. Especially in busy scenarios, this problem becomes particularly prominent. For example, in a business office, staff need to quickly obtain key information provided by the robot to handle urgent matters; or at a medical emergency site, medical staff are racing against time and urgently need the robot to provide relevant data or guidance accurately and concisely. However, due to the large amount of irrelevant content mixed in the robot's voice output, users have to spend more time screening and waiting for key information, resulting in low interaction efficiency.
[0004] Therefore, there is an urgent need for a technology that can optimize the voice interaction content of robots and improve the efficiency of voice interaction to solve this common problem existing in current robot voice interaction and meet the needs of users to obtain information efficiently in different scenarios, especially in busy scenarios. Summary of the Invention
[0005] In response to this, the present invention provides a scene-based robot interaction control method, system, electronic device, computer storage medium, and computer program product to solve the above technical problems.
[0006] The present invention discloses a scenario-based robot interaction control method, which is applied to a robot. The method includes the following steps: scanning to obtain a panoramic image in the application scenario, extracting first personnel feature information from the panoramic image, and analyzing the first personnel feature information to obtain the first busy level of the application scenario; photographing to obtain a wide-angle image in a preset interaction direction area, extracting second personnel feature information from the wide-angle image, and analyzing the second personnel feature information to obtain the second busy level of the application scenario; determining the third busy level of the application scenario according to the first busy level and the second busy level, determining the voice interaction simplification degree according to the third busy level, and performing voice interaction with the object in the interaction direction area according to the voice interaction simplification degree.
[0007] In some embodiments, the analyzing the second busy level of the application scenario according to the second personnel feature information includes: decomposing from the second personnel feature information the first expression and speech rate information of the person in interaction, the first number, second expression and speech rate information of the queuing personnel behind the person in interaction; analyzing the fourth busy level of the person in interaction according to the first expression and speech rate information of the person in interaction, and analyzing the first intervention coefficient according to the first number, second expression and speech rate information of the queuing personnel behind the person in interaction; correcting the fourth busy level according to the first intervention coefficient to obtain the second busy level.
[0008] In some embodiments, the analyzing the second busy level of the application scenario according to the second personnel feature information further includes: decomposing from the second personnel feature information the second number, second expression and speech rate information of the onlookers around the robot on the left and right of the person in interaction; analyzing the second intervention coefficient according to the second number, second expression and speech rate information, and the correcting the fourth busy level according to the first intervention coefficient to obtain the second busy level includes: correcting the fourth busy level according to the first intervention coefficient and the second intervention coefficient to obtain the second busy level.
[0009] In some embodiments, analyzing the second busy level of the application scenario based on the second personnel characteristic information further includes: decomposing from the second personnel characteristic information the interaction intensity between the interacting personnel and the onlookers, and analyzing from the first expression and speech rate information and the second expression and speech rate information the overall emotional positive level of the interacting personnel and the onlookers, and matching a third intervention coefficient according to the interaction intensity and the emotional positive level; wherein, the third intervention coefficient is negatively correlated with both the interaction intensity and the emotional positive level; then analyzing a second intervention coefficient according to the second quantity and the second expression and speech rate information, and the method for correcting the fourth busy level according to the first intervention coefficient to obtain the second busy level includes: using the third intervention coefficient to correct the second intervention coefficient to a fourth intervention coefficient, and correcting the fourth busy level according to the first intervention coefficient and the fourth intervention coefficient to obtain the second busy level.
[0010] In some embodiments, determining the third busy level of the application scenario according to the first busy level and the second busy level includes: setting the weighting coefficient of the first busy level as a first coefficient, the weighting coefficient of the second busy level as a second coefficient, and the first coefficient being less than the second coefficient; performing a fusion calculation on the first busy level and the second busy level based on the first coefficient and the second coefficient to obtain the third busy level of the application scenario.
[0011] The present invention also discloses a scenario-based robot interaction control system applied to a robot. The system includes a processor and a memory. The processor is configured to retrieve and run computer code in the memory to implement the following steps: scanning to obtain a panoramic image in the current application scenario, extracting first personnel characteristic information from the panoramic image, and analyzing the first busy level of the application scenario according to the first personnel characteristic information; photographing to obtain a wide-angle image in a preset interaction direction area, extracting second personnel characteristic information from the wide-angle image, and analyzing the second busy level of the application scenario according to the second personnel characteristic information; determining the third busy level of the application scenario according to the first busy level and the second busy level, determining the voice interaction simplicity degree according to the third busy level, and performing voice interaction with the object in the interaction direction area according to the voice interaction simplicity degree.
[0012] In some embodiments, analyzing the second busyness level of the application scenario based on the second personnel characteristic information includes: decomposing from the second personnel characteristic information the first expression and speech rate information of the interacting person, the first quantity, second expression and speech rate information of the queuing people behind the interacting person; analyzing the fourth busyness level of the interacting person based on the first expression and speech rate information of the interacting person, and analyzing a first intervention coefficient based on the first quantity, second expression and speech rate information of the queuing people behind the interacting person; and correcting the fourth busyness level according to the first intervention coefficient to obtain the second busyness level.
[0013] In some embodiments, analyzing the second busyness level of the application scenario based on the second personnel characteristic information further includes: decomposing from the second personnel characteristic information the second quantity, second expression and speech rate information of the onlookers on the left and right of the interacting person and surrounding the robot; analyzing a second intervention coefficient based on the second quantity, second expression and speech rate information; and the step of correcting the fourth busyness level according to the first intervention coefficient to obtain the second busyness level includes: correcting the fourth busyness level according to the first intervention coefficient and the second intervention coefficient to obtain the second busyness level.
[0014] In some embodiments, analyzing the second busyness level of the application scenario based on the second personnel characteristic information further includes: decomposing from the second personnel characteristic information the interaction intensity between the interacting person and the onlookers, and analyzing the overall positive emotion level of the interacting person and the onlookers based on the first expression and speech rate information and the second expression and speech rate information, and matching a third intervention coefficient according to the interaction intensity and the positive emotion level; wherein the third intervention coefficient is negatively correlated with both the interaction intensity and the positive emotion level; and the step of analyzing a second intervention coefficient based on the second quantity, second expression and speech rate information, and the step of correcting the fourth busyness level according to the first intervention coefficient to obtain the second busyness level includes: correcting the second intervention coefficient to a fourth intervention coefficient using the third intervention coefficient, and correcting the fourth busyness level according to the first intervention coefficient and the fourth intervention coefficient to obtain the second busyness level.
[0015] In some embodiments, determining the third busyness level of the application scenario based on the first busyness level and the second busyness level includes: setting the weighting coefficient of the first busyness level as a first coefficient, the weighting coefficient of the second busyness level as a second coefficient, where the first coefficient is less than the second coefficient; and performing a fusion calculation on the first busyness level and the second busyness level based on the first coefficient and the second coefficient to obtain the third busyness level of the application scenario.
[0016] The present invention also discloses an electronic device, applied to a robot, including: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, where the processor executes the computer program to implement the method as described in any one of the preceding items.
[0017] The present invention also discloses a computer storage medium, applied to a robot, where the computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method as described in any one of the preceding items.
[0018] The present invention also discloses a computer program product, applied to a robot, which, when running on a terminal, causes the terminal to execute to implement the method as described in any one of the preceding items.
[0019] The beneficial effects of the present invention are as follows: By comprehensively analyzing the panoramic image and the wide-angle image of the interaction direction area to determine the busyness level of the scene and adjusting the speech interaction conciseness accordingly, the robot can output more concise and direct speech content in a busy scene, avoiding adding too many irrelevant tone words and greetings, so that users can obtain key information faster, greatly improving the efficiency of speech interaction; and, in a busy scene, users no longer need to spend extra time filtering key information from long speeches, reducing the impatience caused by inefficient interaction, making the entire interaction process smoother and more comfortable, and significantly improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 It is a schematic flowchart of a method for robot interaction control based on a scenario disclosed in an embodiment of the present invention.
[0022] Figure 2It is a schematic diagram of the scenario of the interacting personnel, queuing personnel, and onlookers disclosed in the embodiments of the present invention. Detailed implementation manners
[0023] The following specific embodiments illustrate the implementation manners of the present application. Those skilled in the art can easily understand the other advantages and effects of the present application from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0024] In addition, the technical features involved in different implementation manners of the present application described below can be combined with each other as long as they do not conflict with each other.
[0025] As Figure 1 shown, the embodiments of the present invention disclose a scenario-based robot interaction control method applied to a robot. The method includes the following steps: S01, scanning to obtain a panoramic image in the application scenario, extracting first personnel feature information from the panoramic image, and analyzing the first busy level of the application scenario according to the first personnel feature information.
[0026] The robot scans the panoramic image of the application scenario by, for example, rotating its body or rotating the camera. The panoramic image can comprehensively present the overall situation of the scenario. The first personnel feature information is extracted from the panoramic image, and these information include the number of people, distribution, action states, etc. in the application scenario.
[0027] According to the first personnel feature information, the first busy level of the application scenario can be analyzed. The first busy level is the global busy level of the application scenario. For example, when the people are dense, moving in a hurry, there is a queue and the queue length is long, the first busy level is high; when the people are scarce, moving at a normal speed, there is no queue or the queue length is short, the first busy level is low.
[0028] The first busy level can be obtained through a pre-constructed level classification model. The level classification model is, for example, a model constructed based on decision tree algorithms, SVM algorithms, deep learning algorithms (such as CNN, RNN, Transformer, etc.).
[0029] S02, photographing to obtain a wide-angle image in the preset interaction direction area, extracting second personnel feature information from the wide-angle image, and analyzing the second busy level of the application scenario according to the second personnel feature information.
[0030] The camera equipped on the robot can be a wide-angle camera. During normal interaction, it captures images of a specified area, namely the interaction direction area, which refers to the area directly in front of the robot. The person who needs to interact is located in this area, facing the robot directly. Among them, the larger the viewing angle of the wide-angle camera, the better. For example, a viewing angle of 94 - 118 degrees, which is an ultra-wide angle, is beneficial for capturing all the people around the robot.
[0031] Through image recognition algorithms, the second person feature information can be extracted from the wide-angle image. The second person feature information includes, but is not limited to, the expressions, speech rates, etc. of the interacting people, and can also include the first number of queuing people, the second number of onlookers, etc. According to this second person feature information, the second busyness level of the application scenario can also be analyzed. This second busyness level is the local busyness level of the application scenario. The second busyness level will be elaborated in detail in the subsequent content.
[0032] Similar to the first busyness level, the second busyness level can also be achieved through a pre-constructed level classification model. Preferably, this level classification model is constructed based on deep learning algorithms (such as CNN, RNN, Transformer, etc.).
[0033] S03. Determine the third busyness level of the application scenario according to the first busyness level and the second busyness level, determine the voice interaction simplicity degree according to the third busyness level, and conduct voice interaction with the objects in the interaction direction area according to the voice interaction simplicity degree.
[0034] For the aforementioned first busyness level and second busyness level obtained, the present invention uses an appropriate method to fuse the two, thereby obtaining a third busyness level that can more accurately reflect the busyness degree of the application scenario. This is a more accurate busyness level obtained by comprehensively considering the global busyness of the application scenario and the local busyness of the robot interaction direction area.
[0035] There is a preset corresponding relationship between the third busyness level and the voice interaction simplicity degree. For example, the higher the third busyness level, the higher the voice interaction simplicity degree, and the output voice will be more concise. At this time, more of the filler words and polite greetings involved in the background technology will be omitted, which can significantly improve the voice interaction efficiency; the lower the third busyness level, the lower the voice interaction simplicity degree, and the output voice will be more detailed. At this time, the filler words and polite greetings involved in the background technology will not be omitted or will be omitted less, which can make the voice interaction atmosphere more natural and cordial. Among them, the filler words, polite greetings, etc. are pre-divided into different priorities, and the simplification is carried out according to the priority level. For example, first (corresponding to the situation where the voice interaction simplicity degree is low) simplify the words with low priority, and then (corresponding to the situation where the voice interaction simplicity degree is high) simplify the words with high priority.
[0036] The present invention determines the busyness level of a scene by comprehensively analyzing a panoramic image and a wide-angle image of an interaction direction area, and adjusts the speech interaction conciseness accordingly. The robot can output more concise and direct speech content in a busy scene, avoiding adding too many irrelevant filler words and greetings, so that users can obtain key information faster, greatly improving the efficiency of speech interaction; and, in a busy scene, users no longer need to spend extra time filtering key information from long speeches, reducing impatience caused by inefficient interaction, making the entire interaction process smoother and more comfortable, and significantly improving the user experience.
[0037] It should be noted that the robot of the present invention can be applied to various scenarios such as business offices (such as bank halls, government service halls), production workshops, etc.
[0038] In some embodiments, analyzing the second busyness level of the application scenario according to the second personnel characteristic information includes: decomposing from the second personnel characteristic information the first expression and speech rate information of the person being interacted with, the first number of queuing people behind the person being interacted with, the second expression and speech rate information; analyzing the fourth busyness level of the person being interacted with according to the first expression and speech rate information of the person being interacted with, and analyzing the first intervention coefficient according to the first number, the second expression and speech rate information of the queuing people behind the person being interacted with; correcting the fourth busyness level according to the first intervention coefficient to obtain the second busyness level.
[0039] In the embodiments of the present invention, by analyzing the personnel density, the haste of people's actions, and the queuing situation in the application scenario, the first busyness level reflecting the global busyness degree of the application scenario can be obtained. The change of this first busyness level is slow, so its confidence level is very high. However, the complexity of the personnel characteristic information in the local area of the robot interaction direction area is significantly higher and more variable. Therefore, the present invention sets two steps for analyzing and determining the second busyness level, specifically as follows: decompose two types of key information from the second personnel characteristic information, namely the first expression and speech rate information of the person being interacted with, and the first number, the second expression and speech rate information of the queuing people behind the person being interacted with.
[0040] The facial expression and speaking speed of the person interacting with the robot can reflect their current state and busyness level. For example, an anxious expression and a fast speech rate may indicate that this person is in a relatively busy and anxious state; while a relaxed expression and a slow speech rate may indicate relatively relaxed. The corresponding fourth busyness level can be analyzed and obtained by using the above-mentioned level classification model constructed based on the deep learning algorithm.
[0041] Meanwhile, the first quantity of queuing persons (such as the entity-filled little figures in Figure 2 ) behind the interacting person can also reflect the busyness level of the current interaction scenario. The more the number of persons waiting in line for interaction, the busier this local area is; also, the expressions and speech rates of the queuing persons can equally reflect their states. For example, if the queuing persons look anxious and have a fast speech rate, it indicates that everyone hopes to complete the interaction as soon as possible, and it also means that this local area is relatively busy. Therefore, by analyzing the quantity, expressions, and speech rates of the queuing persons behind the interacting person, the busyness level of this local area can be indirectly determined.
[0042] Based on the above, the present invention first determines the busyness level of the interacting person according to the first expression and speech rate information of the interacting person, obtaining the fourth busyness level. For example, if the interacting person looks nervous and has an extremely fast speech rate, it may be determined that the fourth busyness level is high; if the expression is calm and the speech rate is moderate, it is determined that the fourth busyness level is low. Then, based on the first quantity, second expression, and speech rate information of the queuing persons behind the interacting person, the first intervention coefficient is analyzed. When the number of queuing persons is large, the expressions are anxious, and the speech rate is fast, the corresponding first intervention coefficient is higher, indicating that the persons behind the interacting person also urgently need to conduct voice interaction; on the contrary, if the number of queuing persons is small, the expressions are relaxed, and the speech rate is slow, the corresponding first intervention coefficient is lower. Finally, the first intervention coefficient is used to correct the fourth busyness level, that is, a multiplication operation is performed, so as to obtain the second busyness level.
[0043] In this way, the present invention comprehensively considers the information of the interacting person and the queuing persons, and can more accurately analyze the second busyness level of the application scenario.
[0044] In some embodiments, the analyzing the second busyness level of the application scenario according to the second person feature information further includes: decomposing from the second person feature information the second quantity, second expression, and speech rate information of the onlookers located around the robot on the left and right of the interacting person; analyzing a second intervention coefficient according to the second quantity, the second expression, and the speech rate information, and the correcting the fourth busyness level according to the first intervention coefficient to obtain the second busyness level includes: correcting the fourth busyness level according to the first intervention coefficient and the second intervention coefficient to obtain the second busyness level.
[0045] In the embodiments of the present invention, the obtained second person feature information further includes onlookers located around the robot on the left and right of the interacting person (such as Figure 2Relevant information about the non-entity filled little people in []. Among them, the larger the number of onlookers, the higher the degree of attention of the scene, indicating that the application scenario is busier, and also indicating that the possible first number of people joining the queue and waiting for voice interaction with the robot may increase. The facial expressions and speaking speeds of the onlookers can reflect their states and eagerness for interaction. For example, if the onlookers have eager expressions and fast speaking speeds, it indicates that the probability of these people also hoping to participate in the interaction as soon as possible is greater, and the probability of joining the queue subsequently is correspondingly greater, indirectly reflecting that the application scenario is busier. At this time, the second intervention coefficient is set to be larger; if the expressions are relaxed and the speaking speeds are slow, it indicates that the probability of these people also hoping to participate in the interaction as soon as possible is smaller, and the probability of joining the queue subsequently is correspondingly smaller, indirectly reflecting that the application scenario is less busy. At this time, the second intervention coefficient is set to be smaller.
[0046] Multiply the fourth busyness level by the first intervention coefficient and the second intervention coefficient obtained above simultaneously to obtain the corrected second busyness level.
[0047] By comprehensively considering the situations of queuing people and onlookers and introducing the first intervention coefficient and the second intervention coefficient to correct the fourth busyness level, the present invention can more comprehensively and accurately analyze the second busyness level of the application scenario, so that the robot can adjust the refinement degree of voice interaction according to a more accurate busyness level, improving the voice interaction efficiency and user experience.
[0048] In some embodiments, analyzing the second busyness level of the application scenario according to the second personnel characteristic information further includes: decomposing the interaction intensity between the interacting people and the onlookers from the second personnel characteristic information, and analyzing the overall positive emotion level of the interacting people and the onlookers according to the first expression and speaking speed information and the second expression and speaking speed information, and matching the third intervention coefficient according to the interaction intensity and the positive emotion level; wherein, the third intervention coefficient is negatively correlated with both the interaction intensity and the positive emotion level.
[0049] Then, analyze the second intervention coefficient according to the second quantity, the second expression and speaking speed information. The step of correcting the fourth busyness level according to the first intervention coefficient to obtain the second busyness level includes: using the third intervention coefficient to correct the second intervention coefficient to the fourth intervention coefficient, and correcting the fourth busyness level according to the first intervention coefficient and the fourth intervention coefficient to obtain the second busyness level.
[0050] In the embodiments of the present invention, the foregoing embodiments relate to analyzing the probability of a micro-person subsequently joining a voice interaction queue based on the emotional conditions represented by the micro-person's own expressions and speech rates. However, the micro-person may be in a state of "watching the excitement" rather than having a real voice interaction requirement, and the probability of their subsequent joining the queue is also low. Therefore, it is necessary to consider this factor to ensure that the obtained second intervention coefficient is more accurate, and further ensure the accuracy of the obtained second busy level.
[0051] Specifically, determine the interaction intensity between the person in interaction and the onlookers. For example, this interaction intensity is that the person in interaction has had positive emotional exchanges with multiple onlookers. For example, the person in interaction happily says to the micro-person, "This robot has really good thinking ability." It should be noted that since there may be "acquaintances" of the person in interaction among the micro-people, their interactions with these "acquaintances" cannot accurately represent the probability that these micro-people are "watching the excitement". Therefore, the number of onlookers involved in positive emotional exchanges should be more than a preset number, so as to ensure that the person in interaction has had positive emotional exchanges with multiple strangers, and thus make the conclusion that the micro-person is "watching the excitement" more accurate.
[0052] At the same time, it is also necessary to analyze the overall positive emotional level of all people including the person in interaction and the onlookers. Positive emotions refer to emotions such as happiness and laughter, and their opposite is being silent and having a solemn expression. A corresponding relationship between the interaction intensity, the positive emotional level, and the third intervention coefficient is pre-established, and the third intervention coefficient can be determined by reverse querying this corresponding relationship. Among them, the higher the interaction intensity and the higher the positive emotional level, the greater the probability that these micro-people are "watching the excitement". Correspondingly, the third intervention coefficient is set to be smaller; otherwise, the third intervention coefficient is set to be larger.
[0053] In some embodiments, determining the third busy level of the application scenario according to the first busy level and the second busy level includes: setting the weighting coefficient of the first busy level as the first coefficient, the weighting coefficient of the second busy level as the second coefficient, and the first coefficient being less than the second coefficient; performing a fusion calculation on the first busy level and the second busy level based on the first coefficient and the second coefficient to obtain the third busy level of the application scenario.
[0054] In the embodiments of the present invention, the present invention sets a lower weight for the global busy level of the application scenario and a higher weight for the local busy level, and accordingly performs a fusion process on the foregoing obtained first busy level and second busy level to obtain the final third busy level.
[0055] An embodiment of the present invention also discloses a scenario-based robot interaction control system, which is applied to a robot. The system includes a processor and a memory. The processor is configured to retrieve and run computer code in the memory to implement the following steps: scanning to obtain a panoramic image in the application scenario, extracting first personnel feature information from the panoramic image, and analyzing the first personnel feature information to obtain the first busyness level of the application scenario; capturing a wide-angle image in a preset interaction direction area, extracting second personnel feature information from the wide-angle image, and analyzing the second personnel feature information to obtain the second busyness level of the application scenario; determining a third busyness level of the application scenario according to the first busyness level and the second busyness level, determining a voice interaction refinement degree according to the third busyness level, and performing voice interaction with an object in the interaction direction area according to the voice interaction refinement degree.
[0056] An embodiment of the present invention also discloses an electronic device, which is applied to a robot and includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. The processor executes the computer program to implement the method as described in the foregoing embodiment.
[0057] An embodiment of the present invention also discloses a computer storage medium, which is applied to a robot. The computer storage medium stores a computer program, and the computer program is executed by a processor to implement the method as described in the foregoing embodiment.
[0058] An embodiment of the present invention also discloses a computer program product, which is applied to a robot. When the computer program product runs on a terminal, it causes the terminal to execute to implement the method as described in the foregoing embodiment.
[0059] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0060] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A scene-based robot interactive control method, applied to a robot, characterized in that: The method comprises the following steps: scanning and acquiring a panoramic image in an application scene, extracting first personnel characteristic information from the panoramic image, and analyzing and acquiring a first busyness level of the application scene according to the first personnel characteristic information; photographing and acquiring a wide-angle image in a preset interaction direction area, extracting second personnel characteristic information from the wide-angle image, and analyzing and acquiring a second busyness level of the application scene according to the second personnel characteristic information; determining and acquiring a third busyness level of the application scene according to the first busyness level and the second busyness level, determining and acquiring a voice interaction simplicity according to the third busyness level, and performing voice interaction with an object in the interaction direction area according to the voice interaction simplicity; The second busyness level of the application scenario is obtained according to the analysis of the second person characteristic information, including: decomposing the first expression and speech speed information of the interacting person, and the first number, second expression and speech speed information of the queued persons behind the interacting person from the second person characteristic information; The fourth busyness level of the interacting person is determined according to the first expression and speech speed information of the interacting person; if the interacting person has a nervous expression and speaks very fast, the fourth busyness level is determined to be high; if the interacting person has a calm expression and speaks moderately, the fourth busyness level is determined to be low; A first intervention coefficient is obtained based on the first number, the second expression and the speech speed information of the queued persons behind the interacting person; when the number of queued persons is large, the expressions are anxious and the speech speed is fast, the first intervention coefficient is correspondingly high; when the number of queued persons is small, the expressions are relaxed and the speech speed is slow, the first intervention coefficient is correspondingly low; Correcting the fourth busy level according to the first intervention coefficient to obtain the second busy level; The method further comprises: obtaining a second busy level of the application scenario according to the second personnel characteristic information by decomposing the second personnel characteristic information to obtain a second number, a second expression, and speech speed information of onlookers located on the left and right of the interacting person and surrounding the robot; obtaining a second intervention coefficient according to the second number, the second expression, and the speech speed information; and correcting the fourth busy level according to the first intervention coefficient to obtain the second busy level, including: correcting the fourth busy level according to the first intervention coefficient and the second intervention coefficient to obtain the second busy level; The step of deriving the second busy level of the application scenario according to the analysis of the second personnel characteristic information further includes: decomposing the interaction intensity between the interacting personnel and the onlookers from the second personnel characteristic information, and analyzing the overall emotional positivity level of the interacting personnel and the onlookers according to the first expression and speech rate information and the second expression and speech rate information, and deriving a third intervention coefficient by matching the interaction intensity with the emotional positivity level; wherein the third intervention coefficient is negatively correlated with both the interaction intensity and the emotional positivity level; then deriving the second intervention coefficient according to the analysis of the second quantity and the second expression and speech rate information, and then correcting the fourth busy level according to the first intervention coefficient to obtain the second busy level includes: correcting the second intervention coefficient to a fourth intervention coefficient using the third intervention coefficient, and correcting the fourth busy level according to the first intervention coefficient and the fourth intervention coefficient to obtain the second busy level.
2. The scene-based robot interactive control method according to claim 1, characterized in that: The determining of the third busy level of the application scenario according to the first busy level and the second busy level includes: setting a weighting coefficient of the first busy level as a first coefficient, and a weighting coefficient of the second busy level as a second coefficient, wherein the first coefficient is smaller than the second coefficient; and performing a fusion calculation on the first busy level and the second busy level based on the first coefficient and the second coefficient to obtain the third busy level of the application scenario.
3. A scene-based robot interactive control system, applied to a robot, the system comprising a processor and a memory; characterized in that: The processor is used to retrieve and run the computer code in the memory to implement the following steps: scanning and acquiring a panoramic image in the application scene, extracting first personnel characteristic information from the panoramic image, and analyzing and deriving a first busyness level of the application scene according to the first personnel characteristic information; photographing and acquiring a wide-angle image in a preset interaction direction area, extracting second personnel characteristic information from the wide-angle image, and analyzing and deriving a second busyness level of the application scene according to the second personnel characteristic information; determining and deriving a third busyness level of the application scene according to the first busyness level and the second busyness level, determining and deriving a voice interaction simplicity according to the third busyness level, and performing voice interaction with an object in the interaction direction area according to the voice interaction simplicity; The step of obtaining the second busyness level of the application scenario according to the analysis of the second personnel characteristic information includes: obtaining the first expression and speech speed information of the interacting person, the first number, second expression and speech speed information of the queued persons behind the interacting person from the second personnel characteristic information; The fourth busyness level of the interacting person is determined according to the first expression and speech speed information of the interacting person, wherein if the interacting person has a nervous expression and speaks very fast, the fourth busyness level is determined to be high; if the expression is calm and the speech speed is moderate, the fourth busyness level is determined to be low; A first intervention coefficient is obtained based on the first number, the second expression and the speech speed information of the queued persons behind the interacting person; when the number of queued persons is large, the expressions are anxious and the speech speed is fast, the first intervention coefficient is correspondingly high; when the number of queued persons is small, the expressions are relaxed and the speech speed is slow, the first intervention coefficient is correspondingly low; Correcting the fourth busy level according to the first intervention coefficient to obtain the second busy level; The method further comprises: obtaining a second busy level of the application scenario according to the second personnel characteristic information by decomposing the second personnel characteristic information to obtain a second number, a second expression, and speech speed information of onlookers located on the left and right of the interacting person and surrounding the robot; obtaining a second intervention coefficient according to the second number, the second expression, and the speech speed information; and correcting the fourth busy level according to the first intervention coefficient to obtain the second busy level, including: correcting the fourth busy level according to the first intervention coefficient and the second intervention coefficient to obtain the second busy level; The step of deriving the second busy level of the application scenario according to the analysis of the second personnel characteristic information further includes: decomposing the interaction intensity between the interacting personnel and the onlookers from the second personnel characteristic information, and analyzing the overall emotional positivity level of the interacting personnel and the onlookers according to the first expression and speech rate information and the second expression and speech rate information, and deriving a third intervention coefficient by matching the interaction intensity with the emotional positivity level; wherein the third intervention coefficient is negatively correlated with both the interaction intensity and the emotional positivity level; then deriving the second intervention coefficient according to the analysis of the second quantity and the second expression and speech rate information, and then correcting the fourth busy level according to the first intervention coefficient to obtain the second busy level includes: correcting the second intervention coefficient to a fourth intervention coefficient using the third intervention coefficient, and correcting the fourth busy level according to the first intervention coefficient and the fourth intervention coefficient to obtain the second busy level.
4. An electronic device, applied to a robot, comprising: At least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor executes the computer program to implement the method according to claim 1 or 2.
5. A computer storage medium, applied to a robot, wherein the computer readable storage medium stores a computer program, characterized in that: The computer program is executed by a processor to implement the method according to claim 1 or 2.
6. A computer program product, applied to a robot, characterized in that: When the computer program product runs on a terminal, the terminal is enabled to implement the method according to claim 1 or 2.
Citation Information
Patent Citations
Guide robot and program for guide robot
JP2021018452A
Customized communication assistance
WO2018041714A1