Language and image understanding system based on multiple protocols

By designing a multi-protocol language and image understanding system, the problem of difficulty in obtaining information in public places by deaf and mute people is solved, efficient information transmission and personalized services are achieved, and communication efficiency and comfort of deaf and mute people is improved.

CN119964229APending Publication Date: 2025-05-09周超
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411781546.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2021-11-10
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Deaf and mute people encounter difficulties in obtaining information in public places, and the existing technology is difficult to effectively answer their questions, resulting in low communication efficiency.

Method used

A multi-protocol-based language and image understanding system is designed, including an image collection module, a voice broadcast module and a human body recognition module. Through limb movement speed analysis, disability status analysis and audio-visual splitting units, sign language understanding for deaf and mute people and automatic recognition of various response methods are realized.

Benefits of technology

It improves the efficiency of information acquisition for deaf and mute people in public places, reduces waiting time and enhances the user's comfort through personalized information transmission methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964229A_ABST
    Figure CN119964229A_ABST
Patent Text Reader

Abstract

The invention discloses a language and image understanding system based on multiple protocols, which comprises an image collection module, a voice broadcast module and a human body recognition module, and is characterized in that the human body recognition module comprises a limb movement speed analysis unit, a user disability condition analysis unit and an audio and video splitting unit; the limb movement speed analysis unit is used for measuring the limb movement speed of the user during the period when the limb extends to contact with the robot so as to judge the limb flexibility level of the user, and if the limb flexibility level is high, the playing speed of a picture displayed by a robot display screen is high, so that time is saved, and the subsequent waiting time of the deaf-mute is prevented from being too long; the user disability state analysis unit is used for detecting whether a user has hearing failure or language failure, or both of the hearing failure and the language failure, and the audio and video splitting unit is used for starting a voice function and an image function in different time periods according to a disability analysis report of the user. The method has the advantages of being high in practicability and capable of automatically recognizing sign languages and solving problems in various response modes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sign language, and in particular to a language and image understanding system based on multiple protocols. Background Art

[0002] When deaf-mute people encounter problems in public places, such as unknown terrain and store information, it is inefficient to seek help from pedestrians due to the difficulty of communication. They can seek help from robots. Robots can answer questions based on language and image understanding, effectively improving the comfort of deaf-mute people in this space. Therefore, it is necessary to design a multi-protocol language and image understanding system that is highly practical, automatically recognizes sign language, and solves problems in a variety of response methods. Summary of the invention

[0003] The object of the present invention is to provide a language and image understanding system based on multiple protocols to solve the problems raised in the above background technology.

[0004] In order to solve the above technical problems, the present invention provides the following technical solutions: a multi-protocol-based language and image understanding system, comprising an image collection module, a voice broadcast module and a human body recognition module, characterized in that: the human body recognition module comprises a limb movement speed analysis unit, a user disability status analysis unit and an audio and video separation unit, the limb movement speed analysis unit is used to measure the limb movement speed of the user from the time the limb is extended to the time the limb contacts the robot, so as to determine the user's limb flexibility level, if the limb flexibility level is high, the robot display screen displays the picture at a fast playback speed to save time and avoid the subsequent deaf-mute person waiting for too long, the user disability status analysis unit is used to detect whether the user has a hearing problem or a language problem, or both, and the audio and video separation unit is used to enable the voice function and the image function in time periods according to the user's disability analysis report to reduce unnecessary consumption of the robot battery.

[0005] According to the above technical solution, the user's limb movement speed detection process is as follows:

[0006] The robot is running in a public area. If a user needs help from the robot, he / she should stand in front of the robot and the robot will stop running immediately. The robot detects the current user, scans the current user's height and body shape, and calculates the horizontal distance between the human body and the robot display screen, which is recorded as L. 水平 The calculated distance from the closest point of the human body to the robot display screen to the human hand is recorded as L 垂直 , and then the distance from the human hand to the robot display screen is calculated by the Pythagorean theorem as L 手 ;

[0007] The effective scanning distance of the robot is L 有效距离 , if L 水平Greater than L 有效距离 The robot will issue a voice announcement to remind the user to approach. 水平 Less than or equal to L 有效距离 The robot calculates the user's limb movement speed V 手 , V 手 =L 手 / T 接触时间 , where T 接触时间 The rated user movement speed is set to V, which is the time from when the robot stops running to when the user touches the robot display screen. 额 , V 额 It is divided into six levels, V1-V6. V1 means the user's limbs move at the slowest speed, and V6 means the user's limbs move at the fastest speed. 手 With V 额 By comparing and determining the corresponding voice broadcast speed levels and image playback speed levels, the voice broadcast speed levels are set to A1-A6, A1 indicates the slowest voice broadcast speed, A6 indicates the fastest voice broadcast speed, and the image playback speed levels are set to B1-B6, B1 indicates the slowest image playback speed, and B6 indicates the fastest image playback speed. This achieves the effect of determining the information reception speed based on the speed of human limb movement, while effectively transmitting information and achieving personalized service customization, making users more comfortable when seeking help.

[0008] According to the above technical solution, the process of determining the degree of disability of the deaf-mute is as follows:

[0009] Set users with hearing impairment as level one disability, recorded as C1 personnel; set users with language impairment as level two disability, recorded as C2 personnel; set users with both hearing and language impairment as level three disability, recorded as C3 personnel. The robot display shows that the robot can accept voice services. If the robot cannot receive the user's information within 3 seconds, it will automatically determine that the user is a language-impaired person. At the same time, the robot will issue a voice broadcast, and can perform sign language operations. Click the display to confirm. If the robot cannot receive the user's information within 3 seconds, it will automatically determine that the user is a hearing-impaired person. The above determination process is carried out simultaneously, with a time of 3 seconds. The judgment process ends and the result is analyzed. If the user can speak, the robot accepts the voice information, and the audio and video separation unit drives the image playback function and the voice playback function at this time. The user can obtain effective help information from the image playback. At the same time, it determines whether the user's language organization is fluent. If the fluency is not up to standard, the user is reminded to use sign language. If the fluency is up to standard, no reminder is required. If the user cannot speak but has hearing impairment, the audio and video separation unit only drives the image playback function and the voice playback function at this time. If the user can neither speak nor has hearing impairment, the audio and video separation unit only drives the image playback function and immediately executes the sign language service function.

[0010] According to the above technical solution, the image collection module includes a gesture recognition unit, a surrounding environment interference elimination unit and a jitter amplitude elimination unit. The gesture recognition unit is used to monitor the user's gesture changes to identify the user's intentions. The surrounding environment interference elimination unit is used to mask the noise and alternative dynamic behaviors around the robot to increase the fluency of communication between the user and the robot. The jitter amplitude elimination unit is used to eliminate slight jitters during gesture changes to increase the accuracy of sign language.

[0011] According to the above technical solution, the voice broadcast module includes a voice receiving unit, a voice playing unit and a lip reading recognition unit. The voice receiving unit is used to receive the user's voice information, the voice playing unit is used to play the set voice information to help the user, and the lip reading recognition unit is used to provide lip reading services for the deaf and mute people who do not know sign language.

[0012] According to the above technical solution, the sign language information interaction process is as follows:

[0013] The robot scans the user's dynamic gestures, matches the real-time dynamic gestures with the gesture records in the data repository, translates the user's sign language meaning, and gives corresponding answers based on the translation content to solve the user's problem. During the sign language translation process, the robot predicts the meaning of the user's sentence and provides ten sentences closest to the user's meaning. The sentences are displayed on the display screen and the user can choose from them, which reduces the time it takes for the user to show sign language. Sign language movements are complex, and the language and image understanding system provides multiple options in the form of predictions, which improves efficiency while reducing information errors in sign language expression and increasing accuracy. The user can choose from ten sentences. The user selects the sentence with the closest meaning from the predicted sentences. If the user successfully selects it, the language and image understanding system will answer the question and resolve the user's doubts. If there is no predicted sentence that satisfies the user among the ten sentences, the user can click to exit and continue to display the sign language. The gesture recognition unit continues to receive sign language information and translates it while receiving it. When the translated sign language information is significantly different from the previous prediction, the sentence is predicted again and ten predicted sentences are provided for the user to select. The prediction is repeated until the prediction is successful. If the prediction is still unsuccessful, the user's sign language information is fully translated and an answer is given based on the complete information.

[0014] According to the above technical solution, the language and image understanding system answers the following process:

[0015] In the process of answering questions, picture answers, voice broadcasts and video displays can be selected. The language and image understanding system makes a selection based on the detection information of the user's disability status analysis unit. For C1 personnel, picture answers and video displays can be provided. The video display has a higher priority than the picture answer. The video display information is specific and easy for users to understand. For C2 personnel, picture answers, voice broadcasts and video displays can be provided. The voice broadcast has a higher priority than the voice broadcast and video display. The voice broadcast is efficient. For C3 personnel, picture answers and video displays can be provided. The video display has a higher priority than the picture answer. The video display is to play the answer information on the display screen using sign language.

[0016] During the user's sign language presentation, if the user does not make a choice during the sentence prediction process, and the predicted sentence stays on the display screen for up to 6 seconds, it is judged that the user has poor literacy skills. During the response process of the language and image understanding system, the response method is adjusted, and picture answers are given priority. Picture answers have less text to avoid errors in user understanding. Each answer method will have 2-3 answer methods: picture answer, voice broadcast and image display. If the priority answer method cannot satisfy the user, the user can manually select the answer method until he is satisfied.

[0017] According to the above technical solution, special situation analysis:

[0018] According to the above process, people with acquired hearing loss have the basis of language ability. In this case, the lip reading recognition mode can be selected on the screen displayed by the robot. The lip reading recognition unit finds the lip part according to the facial information scanned by the human body recognition module, and predicts the information the user wants to express based on the lip dynamics. The accuracy of lip reading recognition is poor, and six predicted sentences are given to narrow the selection range and speed up the user's selection. If none of the predicted sentences are selected, continue to collect lip reading information and predict the sentences again. After two lip reading information predictions, the language and image understanding system recommends that the user use sign language for semantic output. If the sentence prediction is successful, voice broadcast and image display are given at the same time.

[0019] According to the above technical solution, the process of eliminating environment and jitter amplitude is as follows:

[0020] Hand tremors are common among the elderly. In the process of gesture recognition, in addition to the dynamic changes of the hands when displaying sign language, there are also slight dynamic amplitudes of the hands caused by hand tremors. The tremor amplitude elimination unit divides the hand dynamic amplitude into 12 levels, Y1-Y12, with Y1 indicating the smallest hand dynamic amplitude and Y12 indicating the largest hand dynamic amplitude. The dynamic amplitudes of the Y1-Y2 levels are automatically eliminated to reduce the error of the language and image understanding system in recognizing sign language;

[0021] Since this space may be a public space with personnel mobility, the surrounding environment interference elimination unit only accepts the display information directly in front of the robot to ensure the uniqueness of the information source.

[0022] According to the above technical solution, the information storage source of the image display is:

[0023] The sign language information is recorded manually and handed over to the animation production company to produce sign language images in a unified format. The sign language images and sign language logic are stored in the language and image understanding system. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0025] Figure 1 is a system schematic diagram of the present invention; DETAILED DESCRIPTION

[0026] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0027] See also Figure 1 The present invention provides a technical solution: a multi-protocol-based language and image understanding system, comprising an image collection module, a voice broadcast module and a human body recognition module, characterized in that the human body recognition module comprises a limb movement speed analysis unit, a user disability status analysis unit and an audio and video separation unit, the limb movement speed analysis unit is used to measure the limb movement speed of the user from the time the limb is extended to the time the limb contacts the robot, so as to determine the user's limb flexibility level, if the limb flexibility level is high, the robot display screen displays the picture at a fast playback speed to save time and avoid the subsequent deaf-mute person waiting for too long, the user disability status analysis unit is used to detect whether the user has a hearing problem or a language problem, or both, the audio and video separation unit is used to enable the voice function and the image function in time periods according to the user's disability analysis report to reduce unnecessary consumption of the robot battery.

[0028] User limb movement speed detection process:

[0029] The robot is running in a public area. If a user needs help from the robot, he / she should stand in front of the robot and the robot will stop running immediately. The robot detects the current user, scans the current user's height and body shape, and calculates the horizontal distance between the human body and the robot display screen, which is recorded as L.水平 The calculated distance from the closest point of the human body to the robot display screen to the human hand is recorded as L 垂直 , and then the distance from the human hand to the robot display screen is calculated by the Pythagorean theorem as L 手 ;

[0030] The effective scanning distance of the robot is L 有效距离 , if L 水平 Greater than L 有效距离 The robot will issue a voice announcement to remind the user to approach. 水平 Less than or equal to L 有效距离 The robot calculates the user's limb movement speed V 手 , V 手 =L 手 / T 接触时间 , where T 接触时间 The rated user movement speed is set to V, which is the time from when the robot stops running to when the user touches the robot display screen. 额 , V 额 It is divided into six levels, V1-V6. V1 means the user's limbs move at the slowest speed, and V6 means the user's limbs move at the fastest speed. 手 With V 额 By comparing and determining the corresponding voice broadcast speed levels and image playback speed levels, the voice broadcast speed levels are set to A1-A6, A1 indicates the slowest voice broadcast speed, A6 indicates the fastest voice broadcast speed, and the image playback speed levels are set to B1-B6, B1 indicates the slowest image playback speed, and B6 indicates the fastest image playback speed. This achieves the effect of determining the information reception speed based on the speed of human limb movement, while effectively transmitting information and achieving personalized service customization, making users more comfortable when seeking help.

[0031] Process for determining the degree of disability of the deaf-mute:

[0032] Set users with hearing impairment as level one disability, recorded as C1 personnel; set users with language impairment as level two disability, recorded as C2 personnel; set users with both hearing and language impairment as level three disability, recorded as C3 personnel. The robot display shows that the robot can accept voice services. If the robot cannot receive the user's information within 3 seconds, it will automatically determine that the user is a language-impaired person. At the same time, the robot will issue a voice broadcast, and can perform sign language operations. Click the display to confirm. If the robot cannot receive the user's information within 3 seconds, it will automatically determine that the user is a hearing-impaired person. The above determination process is carried out simultaneously, with a time of 3 seconds. The judgment process ends and the result is analyzed. If the user can speak, the robot accepts the voice information, and the audio and video separation unit drives the image playback function and the voice playback function at this time. The user can obtain effective help information from the image playback. At the same time, it determines whether the user's language organization is fluent. If the fluency is not up to standard, the user is reminded to use sign language. If the fluency is up to standard, no reminder is required. If the user cannot speak but has hearing impairment, the audio and video separation unit only drives the image playback function and the voice playback function at this time. If the user can neither speak nor has hearing impairment, the audio and video separation unit only drives the image playback function and immediately executes the sign language service function.

[0033] The image collection module includes a gesture recognition unit, a surrounding environment interference elimination unit and a jitter amplitude elimination unit. The gesture recognition unit is used to monitor the user's gesture changes to identify the user's intentions. The surrounding environment interference elimination unit is used to mask the noise and alternative dynamic behaviors around the robot to increase the fluency of communication between the user and the robot. The jitter amplitude elimination unit is used to eliminate slight jitters during gesture changes to increase the accuracy of sign language.

[0034] The voice broadcast module includes a voice receiving unit, a voice playing unit and a lip reading recognition unit. The voice receiving unit is used to receive the user's voice information, the voice playing unit is used to play the set voice information to help the user, and the lip reading recognition unit is used to provide lip reading services for the deaf and mute people who do not know sign language.

[0035] Sign language information interaction process:

[0036] The robot scans the user's dynamic gestures, matches the real-time dynamic gestures with the gesture records in the data repository, translates the user's sign language meaning, and gives corresponding answers based on the translation content to solve the user's problem. During the sign language translation process, the robot predicts the meaning of the user's sentence and provides ten sentences closest to the user's meaning. The sentences are displayed on the display screen and the user can choose from them, which reduces the time it takes for the user to show sign language. Sign language movements are complex, and the language and image understanding system provides multiple options in the form of predictions, which improves efficiency while reducing information errors in sign language expression and increasing accuracy. The user can choose from ten sentences. The user selects the sentence with the closest meaning from the predicted sentences. If the user successfully selects it, the language and image understanding system will answer the question and resolve the user's doubts. If there is no predicted sentence that satisfies the user among the ten sentences, the user can click to exit and continue to display the sign language. The gesture recognition unit continues to receive sign language information and translates it while receiving it. When the translated sign language information is significantly different from the previous prediction, the sentence is predicted again and ten predicted sentences are provided for the user to select. The prediction is repeated until the prediction is successful. If the prediction is still unsuccessful, the user's sign language information is fully translated and an answer is given based on the complete information.

[0037] Language and image understanding system answer process:

[0038] In the process of answering questions, picture answers, voice broadcasts and video displays can be selected. The language and image understanding system makes a selection based on the detection information of the user's disability status analysis unit. For C1 personnel, picture answers and video displays can be provided. The video display has a higher priority than the picture answer. The video display information is specific and easy for users to understand. For C2 personnel, picture answers, voice broadcasts and video displays can be provided. The voice broadcast has a higher priority than the voice broadcast and video display. The voice broadcast is efficient. For C3 personnel, picture answers and video displays can be provided. The video display has a higher priority than the picture answer. The video display is to play the answer information on the display screen using sign language.

[0039] During the user's sign language presentation, if the user does not make a choice during the sentence prediction process, and the predicted sentence stays on the display screen for up to 6 seconds, it is judged that the user has poor literacy skills. During the response process of the language and image understanding system, the response method is adjusted, and picture answers are given priority. Picture answers have less text to avoid errors in user understanding. Each answer method will have 2-3 answer methods: picture answer, voice broadcast and image display. If the priority answer method cannot satisfy the user, the user can manually select the answer method until he is satisfied.

[0040] Special situation analysis:

[0041] According to the above process, people with acquired hearing loss have the basis of language ability. In this case, the lip reading recognition mode can be selected on the screen displayed by the robot. The lip reading recognition unit finds the lip part according to the facial information scanned by the human body recognition module, and predicts the information the user wants to express based on the lip dynamics. The accuracy of lip reading recognition is poor, and six predicted sentences are given to narrow the selection range and speed up the user's selection. If none of the predicted sentences are selected, continue to collect lip reading information and predict the sentences again. After two lip reading information predictions, the language and image understanding system recommends that the user use sign language for semantic output. If the sentence prediction is successful, voice broadcast and image display are given at the same time.

[0042] Environmental rejection and jitter amplitude elimination process:

[0043] Hand tremors are common among the elderly. In the process of gesture recognition, in addition to the dynamic changes of the hands when displaying sign language, there are also slight dynamic amplitudes of the hands caused by hand tremors. The tremor amplitude elimination unit divides the hand dynamic amplitude into 12 levels, Y1-Y12, with Y1 indicating the smallest hand dynamic amplitude and Y12 indicating the largest hand dynamic amplitude. The dynamic amplitudes of the Y1-Y2 levels are automatically eliminated to reduce the error of the language and image understanding system in recognizing sign language;

[0044] Since this space may be a public space with personnel mobility, the surrounding environment interference elimination unit only accepts the display information directly in front of the robot to ensure the uniqueness of the information source.

[0045] The information storage source of the image display:

[0046] The sign language information is recorded manually and handed over to the animation production company to produce sign language images in a unified format. The sign language images and sign language logic are stored in the language and image understanding system.

[0047] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0048] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A multi-protocol based language and image understanding system, comprising an image collection module, a voice broadcast module and a human body recognition module, characterized in that: The human body recognition module includes a limb movement speed analysis unit, a user disability status analysis unit and an audio and video separation unit. The limb movement speed analysis unit is used to measure the limb movement speed of the user from the time the limb is extended to the time the limb contacts the robot, so as to determine the user's limb flexibility level. If the limb flexibility level is high, the robot display screen will display images at a faster speed to save time and avoid the subsequent deaf-mute person waiting for too long. The user disability status analysis unit is used to detect whether the user has a hearing problem or a language problem, or both problems. The audio and video separation unit is used to enable the voice function and the image function in different time periods according to the user's disability analysis report to reduce unnecessary consumption of the robot battery. The robot is running in a public area. If a user needs help from the robot, he / she should stand in front of the robot and the robot will stop running immediately. The robot detects the current user, scans the current user's height and body shape, and calculates the horizontal distance between the human body and the robot display screen, which is recorded as L. 水平 The calculated distance from the closest point of the human body to the robot display screen to the human hand is recorded as L 垂直 , and then the distance from the human hand to the robot display screen is calculated by the Pythagorean theorem as L 手 ; The effective scanning distance of the robot is L 有效距离 , if L 水平 Greater than L 有效距离 The robot will issue a voice announcement to remind the user to approach. 水平 Less than or equal to L 有效距离 The robot calculates the user's limb movement speed V 手 , V 手 =L 手 / T 接触时间 , where T 接触时间 The rated user movement speed is set to V, which is the time from when the robot stops running to when the user touches the robot display screen. 额 , V 额 It is divided into six levels, V1-V6. V1 means the user's limbs move at the slowest speed, and V6 means the user's limbs move at the fastest speed. 手 With V 额 Comparison is made to determine the corresponding voice broadcast speed level and image playback speed level, and the voice broadcast speed level is set to A1-A6, A1 represents the slowest voice broadcast speed, A6 represents the fastest voice broadcast speed, and the image playback speed level is set to B1-B6, B1 represents the slowest image playback speed, and B6 represents the fastest image playback speed, so as to achieve the effect of determining the information reception speed according to the speed of human body movement, and achieve personal service customization while effectively transmitting information, so that users feel more comfortable when seeking help; Set users with hearing impairment as level one disability, recorded as C1 personnel; set users with language impairment as level two disability, recorded as C2 personnel; set users with both hearing and language impairment as level three disability, recorded as C3 personnel. The robot display shows that the robot can accept voice services. If the robot cannot receive the user's information within 3 seconds, it will automatically determine that the user is a language-impaired person. At the same time, the robot will issue a voice broadcast, and can perform sign language operations. Click the display to confirm. If the robot cannot receive the user's information within 3 seconds, it will automatically determine that the user is a hearing-impaired person. The above determination process is carried out simultaneously, with a time of 3 seconds. The judgment process ends in a short time, and the result is analyzed. If the user can speak, the robot receives the voice information, and the audio-visual separation unit drives the image playback function and the voice playback function at this time. The user can obtain effective help information from the image playback, and at the same time, it is determined whether the user's language organization is fluent. If the fluency is not up to standard, the user is reminded to use sign language. If the fluency is up to standard, no reminder is required. If the user cannot speak but has hearing impairment, the audio-visual separation unit only drives the image playback function and the voice playback function at this time. If the user can neither speak nor has hearing impairment, the audio-visual separation unit only drives the image playback function at this time, and immediately executes the sign language service function. The image collection module includes a gesture recognition unit, a surrounding environment interference elimination unit and a jitter amplitude elimination unit. The gesture recognition unit is used to monitor the user's gesture changes to identify the user's intentions. The surrounding environment interference elimination unit is used to mask the noise and alternative dynamic behaviors around the robot to increase the communication fluency between the user and the robot. The jitter amplitude elimination unit is used to eliminate slight jitters during gesture changes to increase the accuracy of sign language. The robot scans the user's dynamic gestures, matches the real-time dynamic gestures with the gesture records in the data repository, translates the user's sign language meaning, and gives corresponding answers based on the translation content to solve the user's problem. During the sign language translation process, the robot predicts the meaning of the user's sentence and provides ten sentences closest to the user's meaning. The sentences are displayed on the display screen and the user can choose from them, which reduces the time it takes for the user to show sign language. Sign language movements are complex, and the language and image understanding system provides multiple options in the form of predictions, which improves efficiency while reducing information errors in sign language expression and increasing accuracy. The user can choose from ten sentences. The user selects the sentence with the closest meaning from the predicted sentences. If the user successfully selects it, the language and image understanding system will answer the question and resolve the user's doubts. If there is no predicted sentence that satisfies the user among the ten sentences, the user can click to exit and continue to display the sign language. The gesture recognition unit continues to receive sign language information and translates it while receiving it. When the translated sign language information is significantly different from the previous prediction, the sentence is predicted again and ten predicted sentences are provided for the user to select. The prediction is repeated until the prediction is successful. If the prediction is still unsuccessful, the user's sign language information is fully translated and an answer is given based on the complete information.

2. The multi-protocol based language and image understanding system according to claim 1, characterized in that: Language and image understanding system answer process: In the process of answering questions, picture answers, voice broadcasts and video displays can be selected. The language and image understanding system makes a selection based on the detection information of the user's disability status analysis unit. For C1 personnel, picture answers and video displays can be provided. The video display has a higher priority than the picture answer. The video display information is specific and easy for users to understand. For C2 personnel, picture answers, voice broadcasts and video displays can be provided. The voice broadcast has a higher priority than the voice broadcast and video display. The voice broadcast is efficient. For C3 personnel, picture answers and video displays can be provided. The video display has a higher priority than the picture answer. The video display is to play the answer information on the display screen using sign language. During the user's sign language presentation, if the user does not make a choice during the sentence prediction process, and the predicted sentence stays on the display screen for up to 6 seconds, it is judged that the user has poor literacy skills. During the response process of the language and image understanding system, the response method is adjusted, and picture answers are given priority. Picture answers have less text to avoid errors in user understanding. Each answer method will have 2-3 answer methods: picture answer, voice broadcast and image display. If the priority answer method cannot satisfy the user, the user can manually select the answer method until he is satisfied.

3. The multi-protocol based language and image understanding system according to claim 1, characterized in that: Environmental rejection and jitter amplitude elimination process: Hand tremors are common among the elderly. In the process of gesture recognition, in addition to the dynamic changes of the hands when displaying sign language, there are also slight dynamic amplitudes of the hands caused by hand tremors. The tremor amplitude elimination unit divides the hand dynamic amplitude into 12 levels, Y1-Y12, with Y1 indicating the smallest hand dynamic amplitude and Y12 indicating the largest hand dynamic amplitude. The dynamic amplitudes of the Y1-Y2 levels are automatically eliminated to reduce the error of the language and image understanding system in recognizing sign language; Since this space may be a public space with personnel mobility, the surrounding environment interference elimination unit only accepts the display information directly in front of the robot to ensure the uniqueness of the information source.

4. The multi-protocol based language and image understanding system according to claim 1, characterized in that: The information storage source of the image display: The sign language information is recorded manually and handed over to the animation production company to produce sign language images in a unified format. The sign language images and sign language logic are stored in the language and image understanding system.