A history celebrity interactive robot based on a vertical field large model
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INNER MONGOLIA NORMAL UNIVERSITY
- Filing Date
- 2026-04-22
- Publication Date
- 2026-08-04
AI Technical Summary
[0004]但是,现有技术中,历史名人交互机器人多依赖单一语音识别方式,易受环境杂音干扰而错误识别观众提问,同时,缺乏完善的反馈机制,难以定位应答准确性不足、输出内容过长等问题,无法为交互优化提供有效支撑,交互体验较差
[0016]与现有技术相比,本发明通过数据采集模块采集观众图像以及采集观众音频,
Smart Images

Figure CN122500752A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interactive robot technology, and in particular to an interactive robot based on a large model of historical figures in a vertical domain. Background Technology
[0002] With the rapid development of artificial intelligence technology, human-computer interaction robots have been widely used in various public scenarios. Among them, interactive robots themed around historical figures have become an important vehicle for cultural dissemination and education due to their combination of popular science and entertainment. However, existing interactive robots based on historical figures mostly rely on pre-set question-and-answer databases for interaction, which suffers from problems such as poor interaction flexibility, inability to accurately capture the audience's interaction intentions, lack of proactive interaction triggering mechanisms, and insufficient interaction feedback optimization capabilities.
[0003] In the prior art, Chinese Patent Publication No. CN107278302A discloses a robot interaction method, which includes the following user information collection steps: obtaining the user's on-site information and calculating user interaction parameters based on the on-site information; when the user interaction parameters meet the requirements, querying and determining the item to be supplemented from the user feature information set, determining the relevant communication scenario information from the communication scenario library based on the item to be supplemented, and actively asking the user a question through voice and / or image based on the communication scenario information related to the item to be supplemented; obtaining the user's voice and / or image feedback information, extracting the relevant content associated with the item to be supplemented from the feedback information, and saving the relevant content to the user feature information set.
[0004] However, in existing technologies, interactive robots for historical figures mostly rely on a single voice recognition method, which is easily affected by ambient noise and may misidentify audience questions. At the same time, they lack a sound feedback mechanism, making it difficult to locate problems such as insufficient accuracy of responses and excessively long output content. This fails to provide effective support for interaction optimization and results in a poor interactive experience. Summary of the Invention
[0005] The purpose of this invention is to provide an interactive robot for historical figures based on a large model in a vertical domain. This robot can rely on multiple methods for identification, is not affected by environmental noise, and has a sound feedback mechanism to provide effective support for interaction optimization and improve the interactive experience.
[0006] This invention provides an interactive robot for historical figures based on a large-scale model in a vertical domain. The robot body integrates a data acquisition module, an analysis module, an output module, and a feedback module. The data acquisition module includes an image acquisition unit disposed at the front end of the robot's head for acquiring images of the audience and an audio acquisition unit for acquiring audio of the audience. The analysis module includes an identification unit and a selection unit. The recognition unit is used to obtain the audience's gaze direction and lip movements based on the audience image, and to identify whether the audience asks a question; The selected unit is used to select an interaction strategy based on whether the audience asks a question, including: Based on the audience images, obtain the audio of the audience Q&A session; Alternatively, calculate the duration of the audience's stay based on the audience's image and determine whether to issue an interactive word based on the audience's gaze direction; The output module is used to output corresponding knowledge information based on the audio during the audience's question-and-answer session. The feedback module is used to obtain audience feedback audio and audience image to determine the interaction results.
[0007] Furthermore, the recognition unit is used to obtain the viewer's gaze direction and lip movements based on the viewer image, wherein, The gaze direction of the audience is identified based on the audience image using a gaze detection algorithm; The image recognition algorithm identifies the audience's lip movements.
[0008] Furthermore, the recognition unit identifies whether the audience asks a question. Obtain the angle of deviation between the viewer's gaze direction and the robot's gaze direction; If the preset conditions are met, the system will identify questions asked by the audience. If the preset conditions are not met, it will be identified that the audience has not asked a question; The preset conditions are that the angle of deviation of the viewer's line of sight is less than a preset deviation threshold, and lip movements are detected.
[0009] Furthermore, the selection unit is used to select an interaction strategy based on whether the audience asks a question, wherein, If a question is identified from an audience member, the audio from the start to the end of the question is obtained based on the audience member's image. If the audience does not ask a question, the system calculates the audience's dwell time based on the audience image and determines whether to issue an interactive word based on the audience's gaze direction.
[0010] Furthermore, the selected unit acquires the audio of the audience's question-asking period based on the audience image. The image recognition algorithm identifies the audience's lip movements, which are then combined with the timestamps recorded by the audio acquisition unit. Determine the start and stop points of lip movements; The audio from the start of the lip movement to the stop of the lip movement is extracted as the audio for the audience's questioning period.
[0011] Furthermore, the selection unit calculates the audience dwell time based on the audience image and determines whether to issue an interactive word based on the audience's gaze direction. The dwell time of the audience is obtained through a target tracking algorithm, where the dwell time is the duration after the audience enters the acquisition range of the image acquisition unit; If an audience member's dwell time exceeds a preset time threshold and their gaze deviation angle is less than a preset deviation threshold, then an interactive word is triggered.
[0012] Furthermore, the output module is used to output corresponding knowledge information based on the audio of the audience's question-and-answer session. Keyword extraction algorithms were used to extract core keywords from the audio during the audience Q&A session; Semantic analysis and intent recognition of keywords are performed using a large-scale model of historical figures in the vertical field. The system retrieves historical figures-related knowledge data stored in the database, performs keyword comparison, and extracts knowledge information corresponding to the audio during the audience's question-and-answer session.
[0013] Furthermore, the feedback module is used to obtain audience feedback audio, wherein the audience feedback audio is the audio of the audience asking a question again.
[0014] Furthermore, the feedback module determines the interaction result. The feedback module receives audience feedback audio transmitted by the audio acquisition unit in real time, and also retrieves the audio from the audience question period; The repetition rate of two questions is calculated using a semantic similarity algorithm; If the repetition between two questions exceeds the preset repetition threshold, it is determined that the knowledge information extracted by the output module does not match the actual needs of the audience. Get the duration of the audience's stay after the output module starts outputting; If the dwell time is less than the preset output time, it is determined that the response content output is too long.
[0015] Furthermore, it also includes a quality statistics module, which records the number of times within a preset time window when the knowledge information extracted by the output module does not match the actual needs of the audience, as well as the number of times the response content is judged to be too long.
[0016] Compared with existing technologies, this invention acquires audience images and audio through a data acquisition module. The recognition unit obtains the audience's gaze direction and lip movements from the audience image to identify whether the audience asks a question. The selection unit selects an interaction strategy based on whether the audience asks a question, including obtaining the audio of the audience's questioning period from the audience image; or calculating the audience's dwell time from the audience image and determining whether to issue an interactive word based on the audience's gaze direction. The output module outputs corresponding knowledge information based on the audio of the audience's questioning period. The feedback module obtains the audience's feedback audio and audience image to determine the interaction result, improving the accuracy and reliability of question recognition, and enhancing the adaptability of human-computer interaction and the audience's experience.
[0017] In particular, the recognition unit of this invention obtains the audience's gaze direction and lip movements based on the audience image to identify whether the audience has asked a question. This effectively avoids the problem of misidentification caused by continuous noise interference in the environment. By simultaneously obtaining the audience's gaze direction and lip movements for dual judgment, the audience's gaze direction can accurately represent whether the audience is paying attention to the robot, and the lip movements can accurately represent whether the audience has initiated a questioning action. The dual judgment logic works together to more effectively, reliably, and accurately identify the audience's questioning behavior, avoiding misjudging environmental noise and conversations of irrelevant people as audience questions. This greatly improves the accuracy and reliability of question recognition and lays the foundation for the accurate selection of subsequent interaction strategies.
[0018] In particular, the selection unit of the present invention selects the interaction strategy based on whether the audience asks a question, which can not only accurately respond to the audience's active questions, but also actively identify potential interaction needs and issue interactive words, breaking the limitations of passive response and improving the initiative of interaction.
[0019] In particular, the feedback module of this invention obtains audience feedback audio and audience images to determine the interaction results. The repetition between the audience feedback audio and the audience question audio can directly characterize the accuracy of the response content. If the repetition is high, it indicates that the knowledge information extracted by the output module does not match the audience's actual needs and has not effectively answered the audience's questions. The audience image can accurately obtain the audience's dwell time after the output module starts outputting. If the audience leaves before the robot finishes outputting, it indicates that the response content output is too long or does not meet the audience's expectations and cannot attract the audience to listen completely. Through the above determination, the interaction effect can be comprehensively and objectively grasped, and the defects in the interaction process can be accurately located. This provides accurate data support for subsequent optimization of output content, adjustment of interaction strategies, and fine-tuning of the vertical domain model, thereby continuously improving the adaptability of human-computer interaction and the audience experience. Attached Figure Description
[0020] Figure 1 This is a structural diagram of the historical figure interactive robot based on a large vertical domain model, according to an embodiment of the present invention. Figure 2 This is a logic decision diagram of the analysis module in an embodiment of the present invention. Detailed Implementation
[0021] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0022] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0023] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.
[0024] Furthermore, it should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0025] like Figure 1 , Figure 2 As shown, this embodiment provides an interactive robot for historical figures based on a large vertical domain model. This interactive robot for historical figures based on a large vertical domain model includes: The robot body integrates a data acquisition module, an analysis module, an output module, and a feedback module. The data acquisition module includes an image acquisition unit located at the front end of the robot's head for acquiring images of the audience and an audio acquisition unit for acquiring audio of the audience.
[0026] In this embodiment, the image acquisition unit uses a 2-megapixel high-definition camera, which is fixedly installed at the front of the robot's head. The acquisition frame rate is set to 30 frames per second, and the acquisition radius is set to 2 meters. It is used to acquire the audience's gaze direction and lip movements in real time, and the image data is transmitted to the analysis module in real time. The audio acquisition unit uses a 4-array omnidirectional microphone with an integrated adaptive noise reduction module, which is installed on both sides of the robot's head. The noise reduction amplitude is set to 40dB. At the same time, it records the acquisition timestamp of each audio segment, and the audio data is transmitted to the analysis module and feedback module in real time.
[0027] The analysis module includes an identification unit and a selection unit. The recognition unit is used to obtain the audience's gaze direction and lip movements based on the audience image, and to identify whether the audience asks a question.
[0028] In this embodiment, when the audience walks into the robot's collection range, looks at the robot, and opens their mouth to ask the robot a question, the recognition unit identifies the audience's gaze deviation angle through a gaze detection algorithm and identifies the audience's lips as having continuous opening and closing movements through an image recognition algorithm.
[0029] In this embodiment, the recognition unit identifies whether an audience member asks a question. Obtain the angle of deviation between the viewer's gaze direction and the robot's gaze direction; If the angle of visual deviation is less than a preset deviation threshold and lip movements are detected, then the audience is identified as making a suggestion.
[0030] In this embodiment, the preset deviation threshold is 15°.
[0031] Understandably, a preset deviation threshold of 15° aligns with human eye-tracking interaction habits. When viewers observe the robot, an eye-tracking deviation of ≤15° can be considered as active attention to the robot, thus avoiding misjudgment due to slight eye-tracking shifts.
[0032] After the selected unit receives the audience's question, it tracks the audience's lip movements using an image recognition algorithm. Combined with the timestamp recorded by the audio acquisition unit, it determines the start and stop points of the lip movements, extracts the audio data within that time period as the audio for the audience's question period, and transmits the audio data to the output module.
[0033] Specifically, the recognition unit of this invention obtains the audience's gaze direction and lip movements based on the audience image to identify whether the audience has asked a question. This effectively avoids misidentification caused by persistent ambient noise. The recognition unit simultaneously obtains the audience's gaze direction and lip movements for dual judgment. The audience's gaze direction accurately indicates whether the audience is paying attention to the robot, while the lip movements accurately indicate whether the audience has initiated a questioning action. The dual judgment logic works together to more effectively, reliably, and accurately identify the audience's questioning behavior, avoiding misjudging ambient noise and conversations of irrelevant people as audience questions. This significantly improves the accuracy and reliability of question recognition, laying the foundation for the accurate selection of subsequent interaction strategies.
[0034] In this embodiment, after the selected unit receives the result that the audience has not asked a question, it calculates the audience's dwell time through a target tracking algorithm. The time is started from when the audience enters the collection range. If the audience is within the collection range, the dwell time is greater than a preset time threshold, and the line of sight deviation angle is less than a preset deviation threshold, the selected unit determines that an interactive word needs to be issued.
[0035] In this embodiment, the preset duration threshold is 3 seconds.
[0036] Understandably, setting a preset duration threshold of 3 seconds can filter out passersby, focus on viewers with potential interaction needs, and prevent the robot from frequently triggering active interaction words, thus affecting the scene experience.
[0037] Specifically, the selection unit of this invention selects an interaction strategy based on whether the audience asks a question, which can not only accurately respond to the audience's active questions, but also actively identify potential interaction needs and issue interactive words, breaking the limitations of passive response and improving the initiative of interaction.
[0038] In this embodiment, the large-scale model for the vertical domain of historical figures is based on the ChatGLM3 general large-scale model and is fine-tuned and trained using labeled data of Tang Dynasty figures; the database adopts graph-based storage to construct nodes and relationships such as Tang Dynasty figures, representative works, and historical events.
[0039] In this embodiment, the output module receives the audio of the audience Q&A session transmitted by the selected unit. First, the audio is preprocessed, and then a keyword extraction algorithm based on a speech coding network model is used to extract core keywords. Subsequently, the keywords are input into a large model of historical figures in the vertical field. The large model performs semantic analysis to clarify the audience's needs. Then, relevant knowledge data in the database is called to compare the keywords and extract the corresponding knowledge information.
[0040] In this embodiment, after the output module outputs knowledge information, the audience asks a question again. The feedback module receives the audio of the second question as the audience feedback audio and retrieves the audio of the previous audience question period. The semantic similarity algorithm is used to calculate the repetition of the two questions. If the repetition of the two questions is greater than a preset repetition threshold, it is determined that the knowledge information extracted by the output module does not match the audience's actual needs.
[0041] In this embodiment, the preset repeatability threshold is 80%.
[0042] Understandably, an 80% similarity score can determine that the audience has not received a satisfactory answer, avoiding omissions due to minor differences in expression. Furthermore, it takes into account the error tolerance of semantic recognition, allowing for slight differences caused by the audience's slip of the tongue or non-standard expression, while avoiding misjudging irrelevant questions as duplicate questions.
[0043] In this embodiment, if the audience's dwell time is less than the preset output time after the output module is started, it is determined that the response content is too long, resulting in the audience not listening to the whole thing.
[0044] In this embodiment, the preset output duration is the knowledge information output duration.
[0045] Specifically, the feedback module of this invention acquires audience feedback audio and audience images to determine the interaction results. The repetition between the audience feedback audio and the audience question audio can directly characterize the accuracy of the response content. If the repetition is high, it indicates that the knowledge information extracted by the output module does not match the audience's actual needs and has not effectively answered the audience's questions. The audience image can accurately obtain the audience's dwell time after the output module starts outputting. If the audience leaves before the robot finishes outputting, it indicates that the response content output is too long or does not meet the audience's expectations and cannot attract the audience to listen completely. Through the above determination, the interaction effect can be comprehensively and objectively grasped, and the defects in the interaction process can be accurately located. This provides accurate data support for subsequent optimization of output content, adjustment of interaction strategies, and fine-tuning of the vertical domain model, thereby continuously improving the adaptability of human-computer interaction and the audience experience.
[0046] In this embodiment, the quality statistics module is used to record the number of times within a preset time window when the knowledge information extracted by the output module does not match the actual needs of the audience, and the number of times the response content is judged to be too long.
[0047] The preset time window is 24 hours to avoid the randomness of short-term data, ensure that the statistical results are of reference value, and avoid optimization lag caused by long-term statistics.
[0048] The modules described in the embodiments of this application can be implemented in software or hardware. These modules can also be located within a processor.
[0049] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using dedicated hardware-based apparatus to perform the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0050] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A historical figure interactive robot based on a large vertical domain model, comprising a robot body, wherein the robot body integrates a data acquisition module, an analysis module, an output module, and a feedback module, characterized in that, The data acquisition module includes an image acquisition unit disposed at the front end of the robot's head for acquiring images of the audience and an audio acquisition unit for acquiring audio of the audience. The analysis module includes an identification unit and a selection unit. The recognition unit is used to obtain the audience's gaze direction and lip movements based on the audience image, and to identify whether the audience asks a question; The selected unit is used to select an interaction strategy based on whether the audience asks a question, including: Based on the audience images, obtain the audio of the audience Q&A session; Alternatively, calculate the duration of the audience's stay based on the audience's image and determine whether to issue an interactive word based on the audience's gaze direction; The output module is used to output corresponding knowledge information based on the audio during the audience's question-and-answer session. The feedback module is used to obtain audience feedback audio and audience image to determine the interaction results.
2. The vertical-domain large model-based historical celebrity interaction robot according to claim 1, wherein The recognition unit is used to obtain the viewer's gaze direction and lip movements based on the viewer image, wherein... The gaze direction of the audience is identified based on the audience image using a gaze detection algorithm; The image recognition algorithm identifies the audience's lip movements.
3. The vertical-domain large model-based historical celebrity interaction robot according to claim 2, wherein, The recognition unit identifies whether the audience asks a question. Obtain the angle of deviation between the viewer's gaze direction and the robot's gaze direction; If the preset conditions are met, the system will identify questions asked by the audience. If the preset conditions are not met, it will be identified that the audience has not asked a question; The preset conditions are that the angle of deviation of the viewer's line of sight is less than a preset deviation threshold, and lip movements are detected.
4. The vertical-domain large model-based historical celebrity interaction robot according to claim 3, wherein, The selected unit is used to select an interaction strategy based on whether the audience asks a question, wherein... If a question is identified from an audience member, the audio from the start to the end of the question is obtained based on the audience member's image. If the audience does not ask a question, the system calculates the audience's dwell time based on the audience image and determines whether to issue an interactive word based on the audience's gaze direction.
5. The vertical-domain large model-based historical celebrity interaction robot according to claim 4, wherein, The selected unit acquires the audio of the audience's question-and-answer session based on the audience image. The image recognition algorithm identifies the audience's lip movements, which are then combined with the timestamps recorded by the audio acquisition unit. Determine the start and stop points of lip movements; The audio from the start of the lip movement to the stop of the lip movement is extracted as the audio for the audience's questioning period.
6. The vertical-domain large model-based historical celebrity interaction robot according to claim 5, wherein, The selected unit calculates the audience dwell time based on the audience image and determines whether to issue an interactive word based on the audience's gaze direction. The dwell time of the audience is obtained through a target tracking algorithm, where the dwell time is the duration after the audience enters the acquisition range of the image acquisition unit; If an audience member's dwell time exceeds a preset time threshold and their gaze deviation angle is less than a preset deviation threshold, then an interactive word is triggered.
7. The vertical-domain large model-based historical celebrity interaction robot of claim 1, wherein, The output module is used to output corresponding knowledge information based on the audio of the audience's question-and-answer session. Keyword extraction algorithms were used to extract core keywords from the audio during the audience Q&A session; Semantic analysis and intent recognition of keywords are performed using a large-scale model of historical figures in the vertical field. The system retrieves historical figures-related knowledge data stored in the database, performs keyword comparison, and extracts knowledge information corresponding to the audio during the audience's question-and-answer session.
8. The vertical-domain large model-based historical celebrity interaction robot of claim 1, wherein, The feedback module is used to obtain audience feedback audio, wherein the audience feedback audio is the audio of the audience asking a question again.
9. The vertical-domain large model-based historical celebrity interaction robot of claim 1, wherein, The feedback module determines the interaction result. The feedback module receives audience feedback audio transmitted by the audio acquisition unit in real time, and also retrieves the audio from the audience question period; The repetition rate of two questions is calculated using a semantic similarity algorithm; If the repetition between two questions exceeds the preset repetition threshold, it is determined that the knowledge information extracted by the output module does not match the actual needs of the audience. Get the duration of the audience's stay after the output module starts outputting; If the dwell time is less than the preset output time, it is determined that the response content output is too long.
10. The interactive robot for historical figures based on a large vertical domain model according to claim 1, characterized in that, It also includes a quality statistics module, which records the number of times within a preset time window when the knowledge information extracted by the output module does not match the actual needs of the audience, as well as the number of times the response content is judged to be too long.