A multi-protocol based language and image understanding system

By designing a multi-protocol-based language and image understanding system, the problem of difficulty in obtaining information in public places by deaf and mute people is solved, and efficient and personalized information transmission and battery energy-saving effects are achieved.

CN114067433BActive Publication Date: 2025-05-30JIANGSU FANGRUAN TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111325893.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-10
Publication Date
2025-05-30
Estimated Expiration
2041-11-10

AI Technical Summary

Technical Problem

Deaf and mute people encounter difficulties in obtaining information in public places, and existing technologies are difficult to effectively solve this problem, especially in terms of unknown terrain and obtaining shop information.

Method used

Design a language and image understanding system based on multi-protocol, including an image collection module, a voice broadcast module and a human body recognition module. The system detects the user's limb movement speed and disability through the human body recognition module, adjusts the activation of voice and image functions, and improves the efficiency of information transmission.

Benefits of technology

It has achieved the efficiency and comfort of obtaining information for deaf and mute people in public places, and has dynamically adjusted voice and image functions to reduce battery consumption and provide personalized services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114067433B_ABST
    Figure CN114067433B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-protocol-based language and image understanding system, including an image collection module, a voice broadcast module, and a human body recognition module. The human body recognition module includes a limb movement speed analysis unit, a user disability condition analysis unit, and an audio-visual splitting unit. The limb movement speed analysis unit is used to measure the limb movement speed of the user during the period from the limb extending to contacting the robot, so as to determine the limb flexibility level of the user. If the limb flexibility level is high, the playback speed of the screen displayed on the robot is fast, thus saving time and avoiding excessive waiting time for the deaf-mute person in the subsequent process. The user disability condition analysis unit is used to detect whether the user has a hearing fault, a language fault, or both. The audio-visual splitting unit is used to enable the voice function and the image function in time periods according to the disability analysis report of the user. The present invention has the characteristics of strong practicability, automatically recognizing sign language, and solving problems with multiple response methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sign language, and specifically to a language and image understanding system based on multiple protocols. Background Art

[0002] When deaf-mutes encounter problems in public places, such as unknown terrain and store information, due to the difficulty of communication for deaf-mutes, it is inefficient to seek help from pedestrians. They can seek help from robots, and the robots can answer questions based on language and image understanding, effectively improving the comfort of deaf-mutes in this space. Therefore, it is necessary to design a language and image understanding system based on multiple protocols with strong practicability, automatic sign language recognition, and various response methods to solve problems. Summary of the Invention

[0003] The purpose of the present invention is to provide a language and image understanding system based on multiple protocols to solve the problems raised in the above background art.

[0004] To solve the above technical problems, the present invention provides the following technical solution: A language and image understanding system based on multiple protocols, including an image collection module, a voice broadcast module, and a human body recognition module. It is characterized in that: the human body recognition module includes a limb movement speed analysis unit, a user disability condition analysis unit, and an audio-visual splitting unit. The limb movement speed analysis unit is used to measure the limb movement speed of the user during the period from the limb stretching out to contacting the robot to determine the limb flexibility level of the user. If the limb flexibility level is high, the display screen of the robot will display the picture at a fast playback speed to save time and avoid the subsequent long waiting time for deaf-mutes. The user disability condition analysis unit is used to detect whether the user has a hearing failure, a language failure, or both. The audio-visual splitting unit is used to enable the voice function and the image function in time periods according to the disability analysis report of the user to reduce unnecessary consumption of the robot battery.

[0005] According to the above technical solution, the detection process of the user's limb movement speed is as follows:

[0006] The robot runs in a public area. If the user needs the help of the robot, the user stands in front of the robot, and the robot immediately stops running. The current user is detected, the height and body shape of the current user are scanned, and the horizontal distance from the human body to the robot display screen is calculated and recorded as L 水平 , the distance from the nearest point of the human body to the robot display screen to the human hand is calculated and recorded as L 垂直 , and then the distance from the human hand to the robot display screen is calculated through the Pythagorean theorem as L 手 ;

[0007] The effective scanning distance of the robot is L 有效距离 , if L 水平Greater than L 有效距离 The robot then issues a voice broadcast to remind the user to approach. If it is L 水平 Less than or equal to L 有效距离 The robot then calculates the limb movement speed V of the user 手 , V 手 = L 手 / T 接触时间 , where T 接触时间 is the time from when the robot stops running to when the user's limb touches the robot's display screen. The rated user movement speed is set as V 额 , and V 额 is divided into six levels V1 - V6. V1 represents the slowest limb movement speed of the user, and V6 represents the fastest limb movement speed of the user. Compare V 手 with V 额 to determine the corresponding voice broadcast speed level and video playback speed level. The voice broadcast speed level is set as A1 - A6, where A1 represents the slowest voice broadcast speed and A6 represents the fastest voice broadcast speed. The video playback speed level is set as B1 - B6, where B1 represents the slowest video playback speed and B6 represents the fastest video playback speed, achieving the effect of determining the information reception speed based on the human body's limb movement speed, achieving personalized service customization while effectively transmitting information, making the user more comfortable when seeking help.

[0008] According to the above technical solution, the process for determining the disability level of a deaf - mute person is as follows:

[0009] Set a user with a single hearing impairment as a first - level disabled person, denoted as a c1 person. Set a user with a single language impairment as a second - level disabled person, denoted as a c2 person. Set a user with both hearing and language impairments as a third - level disabled person, denoted as a c3 person. The robot's display screen shows that this robot can accept voice services. If the robot cannot receive the user's information within 3 seconds, it automatically determines that the user is a language - impaired person. At the same time, a voice broadcast is issued inside the robot, indicating that sign language operations are available. Click on the display screen to confirm. If the robot cannot receive the user's information within 3 seconds, it automatically determines that the user is a hearing - impaired person. The above determination process is carried out simultaneously and ends the determination process in 3 seconds to analyze the result. If the user can speak, the robot receives the voice information. At this time, the audio - video splitting unit drives the video playback function and the voice playback function. The user can obtain effective help information from the video playback. At the same time, it determines whether the user's language organization is fluent. If the fluency does not meet the standard, it reminds the user to use sign language. If the fluency meets the standard, no reminder is required. If the user cannot speak but has normal hearing, the audio - video splitting unit only drives the video playback function and the voice playback function at this time. If the user can neither speak nor hear, the audio - video splitting unit only drives the video playback function at this time and immediately executes the sign language service function.

[0010] According to the above technical solution, the image collection module includes a gesture recognition unit, a surrounding environment interference elimination unit, and a jitter amplitude elimination unit. The gesture recognition unit is used to monitor the gesture changes of the user to identify the user's intention. The surrounding environment interference elimination unit is used to block the noise and other dynamic behaviors around the robot, increasing the communication fluency between the user and the robot. The jitter amplitude elimination unit is used to eliminate the slight jitter during the gesture change process, increasing the accuracy of sign language.

[0011] According to the above technical solution, the voice broadcast module includes a voice receiving unit, a voice playing unit, and a lip language recognition unit. The voice receiving unit is used to receive the voice information of the user. The voice playing unit is used to play the set voice information to assist the user. The lip language recognition unit is used to provide lip language services for deaf-mute people who cannot use sign language.

[0012] According to the above technical solution, the sign language information interaction process is as follows:

[0013] The robot scans the dynamic gestures of the user, matches the real-time dynamic gestures with the gesture records in the data repository, translates the sign language meaning of the user, and makes corresponding responses according to the translation content to solve problems for the user. During the translation of sign language, the system anticipates the meaning of the sentences expressed by the user and provides the ten sentences closest to the meaning expressed by the user, which are displayed on the display screen for the user to select, reducing the time for the user to show sign language. Since the sign language actions are complex, the language and image understanding system provides multiple choices in the form of anticipation, improving the efficiency while reducing the information error of sign language expression and increasing the accuracy rate. The user selects the one with the closest meaning from the ten anticipated sentences. If the user successfully selects, the language and image understanding system answers the question to solve the user's doubts. If there is no anticipated sentence satisfactory to the user among the ten sentences, the user can click to exit and continue to show sign language. The gesture recognition unit continues to receive sign language information and translates the sign language information simultaneously. When the translated sign language information is significantly different from the previous anticipation, sentence anticipation is performed again, providing ten anticipated sentences for the user to select, repeating the anticipation until the anticipation is successful. If the anticipation is never successful, the sign language information of the user is completely translated and answered according to the complete information.

[0014] According to the above technical solution, the answer process of the language and image understanding system is as follows:

[0015] During the process of answering questions, picture answering, voice broadcast, and video display can be selected. The language and image understanding system makes a choice based on the detection information of the user's disability status analysis unit. For c1 personnel, picture answering and video display can be provided. The video display has a higher priority than picture answering. The video display information is specific and easy for users to understand. For c2 personnel, picture answering, voice broadcast, and video display can be provided. The voice broadcast has a higher priority than picture answering and video display. The voice broadcast is efficient. For c3 personnel, picture answering and video display can be provided. The video display has a higher priority than picture answering. The video display is to play the answering information represented by sign language on the display screen;

[0016] During the process of the user's sign language display, if the user does not make a choice during the sentence prediction process and the predicted sentence stays on the display screen for up to 6 seconds, it is determined that the user has weak literacy ability. During the response process of the language and image understanding system, the response method is adjusted, and picture answering is given priority. Picture answering has less text and can avoid errors in user understanding. Each answering method will have 2-3 answering methods including picture answering, voice broadcast, and video display. If the preferred answering method cannot meet the user's needs, the user can manually select the answering method until satisfied.

[0017] According to the above technical solution, special situation analysis:

[0018] People born deaf will all become mute because human language is learned later. They are born unable to hear human language. For this situation, the above process is followed. People who become deaf later have a language ability foundation. For this situation, the lip reading recognition mode can be selected on the screen displayed by the robot. The lip reading recognition unit locates the lip part according to the face information scanned by the human body recognition module, and predicts the information to be expressed by the user according to the lip movement. The accuracy of lip reading recognition is poor, and six predicted sentences are given to narrow the selection range and speed up the user's selection speed. If none of the predicted sentences are selected, the lip language information is continuously collected and the sentence prediction is carried out again. After two lip language information predictions, the language and image understanding system recommends that the user use sign language for semantic output. If the sentence prediction is successful, voice broadcast and video display are given at the same time.

[0019] According to the above technical solution, the environment elimination and jitter amplitude elimination process:

[0020] Senile tremor is common in the elderly. During the gesture recognition process, in addition to the dynamic changes of the hand when showing sign language, there is also a slight dynamic amplitude of the hand caused by senile tremor. The jitter amplitude elimination unit divides the hand dynamic amplitude into 12 levels from Y1 to Y12. Y1 represents the smallest hand dynamic amplitude, and Y12 represents the largest hand dynamic amplitude. The dynamic amplitudes of the Y1-Y2 levels are automatically eliminated to reduce the error of the language and image understanding system in recognizing sign language;

[0021] Since this space may be a public space with personnel mobility, the surrounding environment interference elimination unit only accepts the display information directly in front of the robot to ensure the uniqueness of the information source.

[0022] According to the above technical solution, the information storage source for image display is:

[0023] Sign language information is manually recorded and handed over to an animation production company to produce sign language images in a unified format. Both the sign language images and the sign language logic are stored in the language and image understanding system. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0025] Figure 1 is a schematic diagram of the system of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0027] Please refer to Figure 1 , the present invention provides a technical solution: a language and image understanding system based on multiple protocols, including an image collection module, a voice broadcast module, and a human body recognition module. It is characterized in that: the human body recognition module includes a limb movement speed analysis unit, a user disability condition analysis unit, and an audio-visual splitting unit. The limb movement speed analysis unit is used to measure the limb movement speed of the user during the period from the limb extending to contacting the robot to determine the limb flexibility level of the user. If the limb flexibility level is high, the display screen of the robot will play the picture at a fast speed to save time and avoid the subsequent long waiting time for the deaf-mute. The user disability condition analysis unit is used to detect whether the user has a hearing fault, a language fault, or both. The audio-visual splitting unit is used to enable the voice function and the image function in time periods according to the disability analysis report of the user to reduce the unnecessary consumption of the robot battery.

[0028] User limb movement speed detection process:

[0029] The robot operates in a public area. When a user needs the robot's help, the user stands in front of the robot, and the robot immediately stops running. It detects the current user, scans to obtain the height and body shape of the current user, calculates the horizontal distance from the human body to the robot's display screen and records it as L 水平 It calculates the distance from the closest point of the human body to the robot's display screen to the human hand and records it as L 垂直 Then, through the Pythagorean theorem, it calculates the distance from the human hand to the robot's display screen as L 手 ;

[0030] The effective scanning distance of the robot is L 有效距离 If L 水平 is greater than L 有效距离 the robot issues a voice broadcast to remind the user to get closer. If L 水平 is less than or equal to L 有效距离 the robot calculates the limb movement speed V of the user 手 V 手 = L 手 / T 接触时间 where T 接触时间 is the time from when the robot stops running to when the user's limb touches the robot's display screen. The rated user movement speed is set as V 额 V 额 is divided into six levels V1 - V6. V1 represents the slowest limb movement speed of the user, and V6 represents the fastest limb movement speed of the user. V 手 is compared with V 额 to determine the corresponding voice broadcast speed level and video playback speed level. The voice broadcast speed level is set as A1 - A6, where A1 represents the slowest voice broadcast speed and A6 represents the fastest voice broadcast speed. The video playback speed level is set as B1 - B6, where B1 represents the slowest video playback speed and B6 represents the fastest video playback speed, achieving the effect of determining the information reception speed according to the human limb movement speed, and achieving personalized service customization while effectively transmitting information, making the user more comfortable when asking for help.

[0031] Procedure for determining the disability level of deaf - mutes:

[0032] Set users with single hearing impairment as first-level disabled, denoted as c1 personnel; set users with single language impairment as second-level disabled, denoted as c2 personnel; set users with both hearing and language impairments as third-level disabled, denoted as c3 personnel. The robot display screen shows that this robot can accept voice services. If the robot fails to receive information from the user within 3 seconds, it automatically determines that the user has a language impairment. At the same time, a voice broadcast is issued within the robot, indicating that sign language operations are available. Click on the display screen to confirm. If the robot fails to receive information from the user within 3 seconds, it automatically determines that the user has a hearing impairment. The above determination process is carried out simultaneously, and the determination process ends in 3 seconds to analyze the result. If the user can speak, the robot receives voice information. At this time, the audio-visual splitting unit drives the video playback function and the voice playback function. The user can obtain effective help information from the video playback. At the same time, it is determined whether the user's language organization is fluent. If the fluency does not meet the standard, the user is reminded to use sign language. If the fluency meets the standard, no reminder is required. If the user cannot speak but has normal hearing, the audio-visual splitting unit only drives the video playback function and the voice playback function at this time. If the user can neither speak nor hear, the audio-visual splitting unit only drives the video playback function at this time, and immediately executes the sign language service function.

[0033] The image collection module includes a gesture recognition unit, a peripheral environment interference elimination unit, and a jitter amplitude elimination unit. The gesture recognition unit is used to monitor the gesture changes of the user to identify the user's intention. The peripheral environment interference elimination unit is used to block the noise and other abnormal dynamic behaviors around the robot to increase the communication fluency between the user and the robot. The jitter amplitude elimination unit is used to eliminate the slight jitter during the gesture change process to increase the accuracy of sign language.

[0034] The voice broadcast module includes a voice reception unit, a voice playback unit, and a lip-reading recognition unit. The voice reception unit is used to receive the voice information of the user. The voice playback unit is used to play the set voice information to assist the user. The lip-reading recognition unit is used to provide lip-reading services for deaf-mute people who cannot use sign language.

[0035] Sign language information interaction process:

[0036] The robot scans the user's dynamic gestures, matches the real-time dynamic gestures with the gesture records in the data repository, translates the sign language meaning of the user, and makes corresponding responses according to the translation content to solve problems for the user. During the process of translating sign language, it anticipates the meaning of the sentences expressed by the user and provides the ten sentences closest to the meaning expressed by the user, which are displayed on the display screen for the user to select, reducing the time for the user to show sign language. Since sign language movements are complex, the language and image understanding system provides multiple choices in the form of anticipation, improving efficiency while reducing the information error of sign language expression and increasing the accuracy rate. The user selects the sentence with the closest meaning from the ten anticipated sentences. If the user successfully selects, the language and image understanding system answers the question and solves the user's doubts. If there is no anticipated sentence satisfactory to the user among the ten sentences, the user can click to exit and continue to show sign language. The gesture recognition unit continues to receive sign language information, and while receiving sign language information, it translates sign language information. When the translated sign language information is significantly different from the previous anticipation, sentence anticipation is performed again, providing ten anticipated sentences for the user to select, repeating the anticipation until the anticipation is successful. If the anticipation cannot be successful all the time, the sign language information of the user is completely translated and answered according to the complete information.

[0037] Answer process of the language and image understanding system:

[0038] During the process of answering questions, picture answering, voice broadcast, and video display can be selected. The language and image understanding system makes a choice based on the detection information of the user disability status analysis unit. For c1 personnel, picture answering and video display can be provided, and the video display has a higher priority than picture answering. The video display information is specific and easy for the user to understand. For c2 personnel, picture answering, voice broadcast, and video display can be provided, and the voice broadcast has a higher priority than picture answering and video display. The voice broadcast is efficient. For c3 personnel, picture answering and video display can be provided, and the video display has a higher priority than picture answering. The video display is to play the answering information represented by sign language on the display screen;

[0039] During the process of the user showing sign language, if the user does not make a choice during the sentence anticipation process and the anticipated sentence stays on the display screen for up to 6 seconds, it is determined that the user has weak literacy ability, and the response method is adjusted during the response process of the language and image understanding system, with picture answering taking priority. Picture answering has less text and avoids errors in the user's understanding. Each answering method will have 2 - 3 answering methods including picture answering, voice broadcast, and video display. If the preferred answering method cannot satisfy the user, the user can manually select the answering method until satisfied.

[0040] Analysis of special situations:

[0041] After the above process is completed, people who become deaf later have a foundation in language ability. In this case, the lip-reading recognition mode can be selected on the screen displayed by the robot. The lip-reading recognition unit locates the lip area according to the face information scanned by the human body recognition module, and anticipates the information to be expressed by the user based on the lip movement. The accuracy of lip-reading recognition is poor, and six anticipated sentences are given to narrow the selection range and speed up the user's selection. If none of the anticipated sentences are selected, the lip-reading information is continuously collected and the sentence anticipation is performed again. After two lip-reading information anticipations, the language and image understanding system recommends that the user use sign language for semantic output. If the sentence anticipation is successful, both voice broadcast and video display are given at the same time.

[0042] Environmental exclusion and jitter amplitude elimination process:

[0043] Senile tremors are common in the elderly. During the sign language recognition process, in addition to the dynamic changes of the hands when showing sign language, there are also slight dynamic amplitudes of the hands caused by senile tremors. The jitter amplitude elimination unit divides the hand dynamic amplitude into 12 levels from Y1 to Y12. Y1 represents the smallest hand dynamic amplitude, and Y12 represents the largest hand dynamic amplitude. The dynamic amplitudes of the Y1 - Y2 levels are automatically excluded to reduce the error when the language and image understanding system recognizes sign language.

[0044] Since this space may be a public space with personnel mobility, the surrounding environment interference exclusion unit only accepts the display information directly in front of the robot to ensure the uniqueness of the information source.

[0045] Information storage source of video display:

[0046] The sign language information is manually recorded and handed over to an animation production company to produce sign language videos in a unified format. Both the sign language videos and sign language logic are stored in the language and image understanding system.

[0047] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.

[0048] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A multi-protocol based language and image understanding system, including an image collection module, a voice broadcast module, and a human body recognition module, characterized in that: The human body recognition module includes a limb movement speed analysis unit, a user disability condition analysis unit, and an audio-visual splitting unit. The limb movement speed analysis unit is used to measure the limb movement speed of the user during the period from stretching out the limb to touching the robot, so as to determine the limb flexibility level of the user. If the limb flexibility level is high, the robot display screen will display the video playback speed fast to save time and avoid the subsequent long waiting time for the deaf-mute. The user disability condition analysis unit is used to detect whether the user has a hearing fault, a language fault, or both. The audio-visual splitting unit is used to enable the voice function and the image function in time periods according to the user's disability analysis report to reduce unnecessary consumption of the robot battery; The robot operates in a public area. When the user needs the robot's assistance, the user stands in front of the robot, and the robot immediately stops running. It detects the current user, scans to obtain the height and body shape of the current user, and calculates the horizontal distance from the human body to the robot's display screen, denoted as L 水平 , calculates the distance from the nearest point of the human body to the robot's display screen to the human hand, denoted as L 垂直 , and then calculates the distance from the human hand to the robot's display screen as L through the Pythagorean theorem 手 ; The effective scanning distance of the robot is L 有效距离 , if it is L 水平 greater than L 有效距离 then the robot emits a voice broadcast to remind the user to get closer. If it is L 水平 less than or equal to L 有效距离 then the robot calculates the limb movement speed V of the user 手 , V 手 = L 手 / T 接触时间 , where T 接触时间 is the time from when the robot stops running to when the user's limb touches the robot's display screen. The rated user movement speed is set as V 额 , and V 额 is divided into six levels V1 - V6. V1 indicates the slowest limb movement speed of the user, and V6 indicates the fastest limb movement speed of the user. Compare V 手 with V 额 to determine the corresponding voice broadcast speed level and video playback speed level. The voice broadcast speed level is set as A1 - A6, where A1 indicates the slowest voice broadcast speed and A6 indicates the fastest voice broadcast speed. The video playback speed level is set as B1 - B6, where B1 indicates the slowest video playback speed and B6 indicates the fastest video playback speed, achieving the effect of determining the information reception speed according to the human limb movement speed, realizing personalized service customization while effectively transmitting information, making the user more comfortable when seeking help; Set a user with a single hearing impairment as a first-level disability, denoted as c1 personnel. Set a user with a single language impairment as a second-level disability, denoted as c2 personnel. Set a user with both hearing and language impairments as a third-level disability, denoted as c3 personnel. The robot display screen shows that this robot can accept voice services. If the robot cannot receive the user's information within 3 seconds, it will automatically determine that the user is a language-impaired person. At the same time, a voice broadcast will be issued inside the robot, and sign language operations can be performed. Click on the display screen to confirm. If the robot cannot receive the user's information within 3 seconds, it will automatically determine that the user is a hearing-impaired person. The above determination process is carried out simultaneously and ends the determination process in 3 seconds to analyze the result. If the user can speak, the robot will receive voice information. At this time, the audio-visual splitting unit drives the video playback function and the voice playback function. The user can obtain effective help information from the video playback, and at the same time determine whether the user's language organization is fluent. If the fluency does not meet the standard, the user will be reminded to use sign language. If the fluency meets the standard, there is no need to remind. If the user cannot speak but has normal hearing, the audio-visual splitting unit only drives the video playback function and the voice playback function at this time. If the user can neither speak nor has hearing impairment, the audio-visual splitting unit only drives the video playback function at this time and immediately executes the sign language service function.

2. A multi-protocol based language and image understanding system according to claim 1, characterized in that: The image collection module includes a gesture recognition unit, a surrounding environment interference elimination unit, and a jitter amplitude elimination unit. The gesture recognition unit is used to monitor the gesture changes of the user to identify the user's intention. The surrounding environment interference elimination unit is used to shield the noise and other dynamic behaviors around the robot to increase the communication fluency between the user and the robot. The jitter amplitude elimination unit is used to eliminate the slight jitter during the gesture change process to increase the accuracy of sign language.

3. A multi-protocol based language and image understanding system according to claim 2, characterized in that: The voice broadcast module includes a voice receiving unit, a voice playing unit, and a lip language recognition unit. The voice receiving unit is used to receive the voice information of the user. The voice playing unit is used to play the set voice information to assist the user. The lip language recognition unit is used to provide lip language services for deaf-mute people who cannot sign language.

4. A multi-protocol based language and image understanding system according to claim 3, characterized in that: Sign language information interaction process: The robot scans the dynamic gestures of the user, matches the real-time dynamic gestures with the gesture records in the data repository, translates the meaning of the user's sign language, and makes corresponding answers according to the translation content to solve problems for the user. During the translation of sign language, the meaning of the sentences expressed by the user is predicted, and the ten sentences closest to the meaning expressed by the user are provided and displayed on the display screen for the user to select, reducing the time for the user to show sign language. Since sign language movements are complex, the multi-selection provided by the language and image understanding system in the form of prediction can improve efficiency while reducing the error of sign language expression information, increasing the accuracy rate. The user selects the one with the closest meaning from the ten predicted sentences. If the user successfully selects, the language and image understanding system answers the question to solve the user's doubts. If there is no predicted sentence satisfactory to the user among the ten sentences, the user can click to exit and continue to show sign language. The gesture recognition unit continues to receive sign language information and translate the sign language information at the same time. When the translated sign language information is significantly different from the previous prediction, sentence prediction is performed again, and ten predicted sentences are provided for the user to select. The prediction is repeated until the prediction is successful. If the prediction cannot be successful all the time, the sign language information of the user is completely translated and answered according to the complete information.

5. A multi-protocol based language and image understanding system according to claim 4, characterized in that: Language and image understanding system answer process: During the process of answering questions, picture answering, voice broadcast, and video display can be selected. The language and image understanding system makes a choice based on the detection information of the user disability status analysis unit. For c1 personnel, picture answering and video display can be provided, and the video display has a higher priority than picture answering. The video display information is specific and easy for the user to understand. For c2 personnel, picture answering, voice broadcast, and video display can be provided, and the voice broadcast has a higher priority than picture answering and video display. The voice broadcast is efficient. For c3 personnel, picture answering and video display can be provided, and the video display has a higher priority than picture answering. The video display is to play the answering information expressed in sign language on the display screen; During the user's sign language display process, if the user does not make a selection during the sentence prediction process and the predicted sentence stays on the display screen for up to 6 seconds, it is determined that the user has weak literacy ability. During the response process of the language and image understanding system, the response method is adjusted, and picture answers are given priority. Picture answers have less text, avoiding errors in the user's understanding. Each response method will have 2-3 response methods including picture answers, voice broadcasts, and video displays. If the preferred response method cannot satisfy the user, the user can manually select the response method until satisfied.

6. A language and image understanding system based on multiple protocols according to claim 5, characterized in that: Analysis of special cases: People who become deaf later in life have a foundation in language ability. For this situation, the lip-reading recognition mode can be selected on the screen displayed by the robot. The lip-reading recognition unit locates the lip area according to the face information scanned by the human body recognition module, and predicts the information to be expressed by the user based on the lip movement. The accuracy of lip-reading recognition is poor, and six predicted sentences are given to narrow the selection range and speed up the user's selection speed. If none of the predicted sentences are selected, lip-reading information will be collected again and sentence prediction will be performed again. After two lip-reading information predictions, the language and image understanding system recommends that the user use sign language for semantic output. If the sentence prediction is successful, voice broadcasts and video displays will be given simultaneously.

7. A language and image understanding system based on multiple protocols according to claim 6, characterized in that: Process of environmental elimination and jitter amplitude elimination: Senile tremors are common in the elderly. During the gesture recognition process, in addition to the dynamic changes of the hand during sign language display, there are also slight dynamic amplitudes of the hand caused by senile tremors. The jitter amplitude elimination unit divides the hand dynamic amplitude into a total of 12 levels from Y1 to Y12. Y1 represents the smallest hand dynamic amplitude, and Y12 represents the largest hand dynamic amplitude. The dynamic amplitudes of the Y1-Y2 levels are automatically eliminated to reduce the error when the language and image understanding system recognizes sign language; Since this space may be a public space with personnel mobility, the surrounding environment interference elimination unit only accepts the display information directly in front of the robot to ensure the uniqueness of the information source.

8. A language and image understanding system based on multiple protocols according to claim 7, characterized in that: Information storage source of video display: Sign language information is recorded manually and handed over to an animation production company to produce sign language videos in a unified format. Both the sign language videos and sign language logic are stored in the language and image understanding system.

Citation Information

Patent Citations

  • Auxiliary semantic recognition method based on gesture recognition

    CN111144367A