Control method and device based on deaf-mute travel auxiliary vision glasses and medium

Through the optical waveguide display module and front-facing image collector of the travel assisted visual glasses of the deaf and mute, combined with gesture recognition and semantic integration model, the accuracy and interactivity problems of sign language translation equipment in complex environments is solved, efficient and accurate sign language translation and environmental perception are achieved, and the safety and communication experience of deaf and mute travel are improved.

CN120335601APending Publication Date: 2025-07-18SHANTOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510389510.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing sensor-based sign language translation devices are susceptible to external environmental factors, resulting in inaccurate data collection and poor results in identifying complex gestures, which affects the interactivity and experience between the user and the talker.

Method used

The travel-assisted visual glasses of the deaf and mute people using optical waveguide display module and front-facing image collector combine gesture recognition model and semantic integration big model to capture sign language actions and environmental information in real time. Through cache ring queues and semantic integration technology, accurate output statements are generated, and two-way interaction is achieved with the voice recognition module.

Benefits of technology

It improves the accuracy and user experience of sign language recognition and translation, reduces the pressure on computing resources, and enhances the communication ability and security of deaf and mute people in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335601A_ABST
    Figure CN120335601A_ABST
Patent Text Reader

Abstract

The invention provides a control method and device based on deaf-mute travel auxiliary vision glasses and a medium, and belongs to the technical field of auxiliary vision. The method comprises the following steps: acquiring current first image information, and detecting whether a gesture action exists or not according to a set gesture recognition model; if yes, the current first image information is recognized according to the set gesture recognition model, single isolated word information is extracted and obtained, and the isolated word information is added to the cache annular queue one by one according to the time sequence. And when the quantity of the isolated word information is greater than the set quantity threshold value, correspondingly reading the context information stored in the current storage text. And processing the isolated word information and the context information according to the set semantic integration large model, capturing the context information, carrying out semantic integration on the isolated word information according to the context information and the context information, and displaying an obtained output statement. The data acquisition accuracy, the sign language recognition translation accuracy and the user experience can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of assisted vision technology, and in particular to a control method, device and medium for a visual glasses for the travel assistance of deaf-mutes. Background Art

[0002] Current sensor-based sign language translation devices are vulnerable to external environmental factors, resulting in inaccurate data collection, increased usage costs and maintenance complexity. Conventional visual translation devices have limited detection ranges and accuracies, and can only recognize simple and fixed sign language actions. The recognition and translation effects for complex gestures are not good, and some devices based on traditional image recognition algorithms are prone to recognition errors or results where the translated content fails to convey the exact meaning in complex backgrounds, changing lighting conditions, and hand occlusion situations, resulting in poor interactivity between the user and the converser and affecting the user experience. Summary of the Invention

[0003] The main purpose of the embodiments of this application is to propose a control method, device and medium for a visual glasses for the travel assistance of deaf-mutes, so as to improve the data collection accuracy rate, the sign language recognition and translation accuracy, and the user experience.

[0004] To achieve the above object, on the one hand, an embodiment of this application proposes a control method for a visual glasses for the travel assistance of deaf-mutes, and the method includes:

[0005] The visual glasses include a waveguide display module, a front image collector, a frame and temple arms;

[0006] One side of the frame is provided with a fixing groove, the front image collector is arranged in the fixing groove, the front image collector is used to obtain first image information, and the waveguide display module is arranged at the front of the frame;

[0007] The control method includes:

[0008] Obtain the current first image information, and detect whether there is a sign language action in the current first image information according to the set sign language recognition model;

[0009] If so, recognize the current first image information according to the set sign language recognition model, extract single isolated word information, and add the isolated word information to the cache circular queue one by one according to the time sequence;

[0010] When the number of the isolated word information in the cache circular queue is greater than the set quantity threshold, read the context information stored in the current stored text according to the current isolated word information;

[0011] Integrate the large model according to the set semantics, process the isolated word information and the context information, capture the context information, and semantically integrate the isolated word information according to the context information and the context information to obtain an integrated output statement;

[0012] Display the output statement in text form in the optical waveguide display module.

[0013] Furthermore, according to the set gesture recognition model, recognize the current first image information, and extract single isolated word information, including:

[0014] Perform recognition and inference on the current first image information through the set gesture recognition model to recognize sign language isolated words and obtain the confidence of the sign language isolated words;

[0015] When the confidence is within the set threshold range, then within the set time period, recognize the current first image information multiple times, obtain the corresponding confidence and record the current recognition times until the current recognition times reach the set number threshold;

[0016] When the mean of the corresponding confidence is greater than the set confidence threshold, then use the sign language isolated word as the single isolated word information.

[0017] Furthermore, the integration process of the output statement includes:

[0018] Use the read isolated word information as the first input feature and the read context information as the second input feature;

[0019] According to the set semantic integration large model, obtain the set prompt word template, and determine the context information according to the first input feature and the set prompt word template;

[0020] Input the context information, the set prompt word template and the second input feature into the set semantic integration large model to semantically integrate the first input feature to obtain an output feature;

[0021] Parse the output feature to obtain the used isolated word information and the output statement containing the used isolated word information.

[0022] Furthermore, the control method further includes:

[0023] Delete the used isolated word information from the cache circular queue;

[0024] Store the output statement in text form in the current stored text to update the context information.

[0025] Further, the vision glasses further include a wide-angle image collector, which is disposed at the end of one of the temple arms. The wide-angle image collector is used to obtain second image information. The control method further includes:

[0026] Obtain the second image information, and detect whether there is vehicle information in the second image information according to the set vehicle detection model;

[0027] If so, process the second image information according to the set vehicle detection model, output the bounding box coordinates of the vehicle information, obtain the structural parameters of the wide-angle image collector, and estimate the position information of the vehicle information according to the structural parameters and the bounding box coordinates;

[0028] According to the set vehicle detection model, compare the bounding box coordinates corresponding to the vehicle information in consecutive frames, and estimate the motion trend of the vehicle information according to the corresponding bounding box coordinates;

[0029] Display the position information and the motion trend in the waveguide display module

[0030] Further, the control method further includes:

[0031] When it is confirmed that there is a gesture action in the current first image information, compare the gesture action with the set voice-playing gesture action;

[0032] When it is confirmed that the gesture action is the set voice-playing gesture action, obtain a voice-playing instruction, and play the output statement according to the voice-playing instruction.

[0033] Further, the vision glasses further include a voice recognition module. The control method further includes:

[0034] When it is confirmed that there is a gesture action in the current first image information, compare the gesture action with the set voice-activation gesture action;

[0035] When it is confirmed that the gesture action is the set voice-activation gesture action, obtain a voice recognition activation instruction, and activate the voice recognition module;

[0036] According to the voice recognition activation instruction, obtain ambient voice information, convert the ambient voice information into text information, and display the text information in the waveguide display module or play the text information.

[0037] Further, when the confidence level is within the set threshold range, it includes:

[0038] When the confidence level is less than the minimum threshold in the set threshold range, the sign language isolated word corresponding to the confidence level is removed.

[0039] To achieve the above object, on the other hand, an embodiment of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above method is implemented.

[0040] To achieve the above object, on the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above method is implemented.

[0041] The embodiments of the present application at least include the following beneficial effects: The present application provides a control method, device and medium for a visual aid for deaf-mute people to travel. This solution sets a front image collector on one side of the spectacle frame to capture sign language actions and perceive the surrounding environment in real time. The arrayed waveguide display technology is used as the display part of the visual glasses, enabling clear observation of the real environment and the output sentences can also be superimposed and displayed in the field of view without blocking the user's line of sight. The set gesture recognition model is used to detect the first image information to complete the detection and recognition of the sign language isolated word of a single gesture, and the set semantic integration large model is used to accurately capture the context information according to the isolated word information and the context information, and combine the context information to perform semantic integration on the isolated word information to obtain an output sentence that can express accurately. Part of the computing pressure is shifted to the large model, reducing the burden on the inference resources of the set gesture recognition model, improving the operation efficiency, reducing the computing resource pressure, having better sign language context understanding ability, being able to generate more reasonable semantic sign language texts, improving the understanding ability of the translation model for sign language translation tasks, improving the translation efficiency and accuracy, and improving the user experience. Description of the Drawings

[0042] Figure 1 is a flowchart of the control method for a visual aid for deaf-mute people to travel provided by an embodiment of the present application;

[0043] Figure 2 is a schematic diagram of the spectacle structure of a visual aid for deaf-mute people to travel provided by an embodiment of the present application;

[0044] Figure 3 is a schematic diagram of the main control module structure of a visual aid for deaf-mute people to travel provided by an embodiment of the present application;

[0045] Figure 4 is a schematic diagram of the display area of a visual aid for deaf-mute people to travel provided by an embodiment of the present application;

[0046] Figure 5It is a schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present application.

[0047] Reference numerals: frame 100, fixing groove 110, temple 200, front image collector 300, wide-angle image collector 400, waveguide display module 500, cabin upper cover 600, cabin box 610, main control cabin 620, cabin fixing bracket 630. Detailed implementation manners

[0048] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application detailed in the appended claims.

[0049] It can be understood that the terms "first", "second", etc. used in the present application can be used in this article to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information. Depending on the context, the words "if", "when" as used herein can be interpreted as "when...", "when...", or "in response to determining".

[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0051] Before elaborating on the embodiments of the present application in detail, some nouns and terms involved in the embodiments of the present application will be described first. The nouns and terms involved in the embodiments of the present application are applicable to the following explanations.

[0052] Augmented Reality (AR) glasses are a type of technology device that combines virtual information with the real world. They enhance the user's perception of the real world by overlaying digital content in the user's field of view.

[0053] Circular queue: uses an array of fixed size, but organizes data in a circular manner. That is, the last element of the array is followed by the first element of the array.

[0054] Reference Figure 2 Figure 2 , an embodiment of the present application provides a visual glasses for the travel assistance of deaf-mutes, which includes: a frame 100, temple arms 200, a front image collector 300, a wide-angle image collector 400, a voice playback module, a voice recognition module, and a waveguide display module 500.

[0055] Among them, the visual glasses are AR glasses.

[0056] A fixing groove 110 is provided on one side of the frame 100 close to one of the temple arms 200. The fixing groove 110 is used to fix the front image collector 300. The collection surface of the front image collector 300 faces the front of the glasses. The front image collector 300 can collect the sign language images in the front to obtain the first image information.

[0057] The wide-angle image collector 400 is installed at the end of one of the temple arms 200. A fixing groove 110 is also provided at the end of one of the temple arms 200 to install the wide-angle image collector 400. The collection surface of the wide-angle image collector 400 faces the rear of the glasses. The wide-angle image collector 400 can collect the environmental images in the rear to obtain the second image.

[0058] By setting the front image collector 300 and the wide-angle image collector 400 in the visual glasses to collect images, the sign language actions can be captured and the environment can be assisted in perception. Depending on the wearable device of the glasses, the user can wear them comfortably. Compared with the sign language translation devices that need to wear specific sensors in the prior art, the sensor devices set on the visual glasses are reduced to the greatest extent, the limitations brought by the redundant sensors are reduced, and it can be better integrated into the daily life of the user.

[0059] A lens circuit board is provided on the temple arm 200. The lens circuit board can perform data interaction with the main control module. The waveguide display module 500 is installed in the frame 100 and serves as the display part of the visual glasses to display text information or output statements. The voice recognition module and the voice playback module are both built into the lens circuit board, that is, located on the temple arm 200. The voice recognition module can collect the environmental voice to obtain the environmental voice information.

[0060] In addition, the waveguide display module 500 adopts a single-layer waveguide design with a thickness of 0.7 mm. The waveguide display module 500 is paired with Micro-LED, which is small in size and light to wear. It makes the visual glasses reach a new height in terms of display effect and user experience, and provides a more natural and comfortable interaction experience for the deaf community.

[0061] By setting up a speech recognition module and a speech playback module, the vision glasses can not only achieve sign language recognition but also have voice interaction, improving the conversation interactivity between the user and the interlocutor, providing a more convenient and efficient communication method for the deaf and mute group, and facilitating the understanding of hearing people. By setting up the waveguide display module 500 as the display part of the vision glasses, the array waveguide display technology can achieve a display effect with high transparency and high resolution. When wearing the vision glasses, the user can clearly see the real environment around, and at the same time, the sign language translation sentences can also be superimposed and displayed in the field of vision in high quality. This high-transparency display method does not block the user's line of sight, ensuring that the user can naturally observe the expressions and movements of the other party when communicating with others, enhancing the authenticity and naturalness of communication.

[0062] Among them, referring to Figure 3 , the shell of the main control module includes: a cabin upper cover 600, a cabin box 610, a main control cabin 620, and a cabin fixing bracket 630. The main control module is installed inside the main control cabin 620. The cabin upper cover 600 is arranged above the cabin box 610, and the cabin upper cover 600 is slidably connected to the cabin box 610, which is convenient for maintaining the main control module. The cabin fixing bracket 630 is provided on the cabin upper cover 600, and the cabin fixing brackets 630 are also provided on both sides of the cabin box 610. The user can wear it through the cabin fixing bracket 630 to maintain stability.

[0063] Figure 1 is an optional flowchart of the control method for the vision glasses for the travel assistance of the deaf and mute provided by the embodiment of the present application. Figure 1 The method in

[0064] S100, obtain the current first image information, and according to the set gesture recognition model, detect whether there is a gesture action in the current first image information.

[0065] S200, if so, then according to the set gesture recognition model, recognize the current first image information, extract the single isolated word information, and add the isolated word information to the cache circular queue one by one according to the time sequence.

[0066] S300, when the number of isolated word information in the cache circular queue is greater than the set quantity threshold, then according to the current isolated word information, read the context information stored in the current storage text correspondingly.

[0067] S400, process the isolated word information and the context information according to the set semantic integration large model, capture the context information, and integrate the semantic information of the isolated word information according to the context information and the context information to obtain the integrated output sentence.

[0068] In S500, the output statement is displayed in text form on the optical waveguide display module.

[0069] In the S100 to S500 illustrated in the embodiments of the present application, by setting a front image collector at the connection between one of the temple arms and the frame, so as to capture sign language actions in real time and sense the surrounding environment, and adopting the array optical waveguide display technology as the visual glasses display part, it is possible to clearly observe the real environment and the output statement can also be superimposed and displayed in the field of view without blocking the user's line of sight. By using the set gesture recognition model to detect the first image information, the detection and recognition of single-gesture sign language isolated words are completed, and the set semantic integration large model is used to accurately capture the context information according to the isolated word information and the context information, and contact the context information to perform semantic integration on the isolated word information, so as to obtain an output statement that can express accurately, shift part of the computational pressure to the large model, reduce the burden of the inference resources of the set gesture recognition model, improve the operation efficiency, reduce the computational resource pressure, have better sign language context understanding ability, can generate more reasonable semantic sign language texts, improve the understanding ability of the translation model for sign language translation tasks, improve the translation efficiency and accuracy, and improve the user experience.

[0070] In S100 of some embodiments, the current first image information is obtained through the front image collector, and the first image information is input as an input feature into the set gesture recognition model, and the set gesture recognition model is used to perform recognition and inference on the first image information to detect whether there is a gesture action in the first image information.

[0071] In S200 of some embodiments, when it is determined that there is a gesture action, the set gesture recognition model is used to perform sign language isolated word recognition on the first image new type to obtain single isolated word information. The isolated word information is saved to the cache circular queue one by one according to the time sequence, providing basic data for subsequent semantic integration.

[0072] The circular queue is used as a data cache and as a temporary data storage structure to ensure that subsequent semantic integration can respond quickly, maintain the orderliness of the data, and reduce the burden on memory resources.

[0073] Among them, the set gesture recognition model can perform recognition and inference on one or two gesture actions to determine the corresponding sign language isolated words. The set gesture recognition model can be the YOLOv11 model.

[0074] In S300 of some embodiments, it is judged whether the number of isolated word information in the cache circular queue is greater than the set quantity threshold.

[0075] If so, obtain the local reading path of the stored text through the current isolated word information, and read the current stored text according to the local reading path to obtain the context information corresponding to the current isolated word information, so as to call the set semantic integration large model subsequently.

[0076] Among them, the stored text is a local file. The set quantity threshold can take the value of one isolated word information, and can also take the value of two isolated word information, or other values. In this embodiment, no specific limit is imposed on the set quantity threshold.

[0077] Exemplarily, determine whether there are two or more isolated word information in the cache circular queue. If there are two or more isolated word information, read the corresponding context information through the current isolated word information, so as to call the set semantic integration large model.

[0078] Confirm whether to obtain the corresponding context information by detecting the quantity of the isolated word information, so as to call the set semantic integration large model, reduce misjudgment, and improve the recognition efficiency.

[0079] In S400 of some embodiments, read the isolated word information and context information in the cache circular queue, and input both the isolated word information with time series and the corresponding context information into the set semantic integration large model.

[0080] Through the set semantic integration large model, capture the context information of the sign language expression, and combine with the context information to perform semantic integration on the isolated word information, supplement the missing words, and obtain the integrated output statement.

[0081] Compared with the prior art solution that uses a CNN convolutional neural network for image feature processing and uses a time series processing network such as an LSTM for sign language recognition and classification, the present application can better improve the inference efficiency through the combination of the set semantic integration large model and the set gesture recognition model, realize the recognition of isolated words of a single gesture, and at the same time integrate the context semantics, can clearly express the context, avoid the phenomenon of failing to convey the intended meaning, shift part of the operation pressure to the set semantic integration large model, reduce the burden on the inference resources, improve the operation efficiency, make full use of the limited large model, have better sign language context understanding ability, and generate more reasonable semantic sign language text.

[0082] In S500 of some embodiments, refer to Figure 4 , through the text parameters of the output statement, determine the set display text area in the optical waveguide display module according to the text parameters, and render and display the output statement in the display text area.

[0083] In some embodiments of the present application, S200 may include but is not limited to S210 to S230:

[0084] S210. Identify and infer the current first image information through the set gesture recognition model, identify the sign language isolated word, and obtain the confidence of the sign language isolated word.

[0085] S220. When the confidence is within the set threshold range, within the set time period, repeatedly identify the current first image information multiple times, obtain the corresponding confidence and record the current number of identifications until the current number of identifications reaches the set number threshold.

[0086] S230. When the mean value of the corresponding confidence is greater than the set confidence threshold, then regard the sign language isolated word as a single isolated word information.

[0087] In an embodiment of S210, when it is determined that there is a gesture action, use the set gesture recognition model to identify and infer the first image information, and identify the sign language isolated word corresponding to the gesture action.

[0088] Calculate the confidence of the sign language isolated word by using the confidence scoring method in the set gesture recognition model.

[0089] In an embodiment of S220, if the confidence is within the set threshold range, then repeatedly identify and infer the current first image information, obtain the confidence of each identification and inference, and record the number of repeated identifications and inferences to obtain the current number of identifications until within the set time period, the current number of identifications reaches the set number threshold.

[0090] That is to say, within the set time period, according to the set gesture recognition model and the first image information, repeatedly identify and infer the set number threshold, and obtain the confidence of each repetition, that is, the corresponding confidence.

[0091] Among them, the set threshold range can take a value of [0.7 - 0.85]. When the confidence is between 0.7 and 0.85, multiple identifications are required.

[0092] In an embodiment of S230, calculate the mean value according to the corresponding confidence, and determine whether the mean value is greater than the set confidence threshold.

[0093] If so, regard the corresponding sign language isolated word as a single isolated word information to be stored in the cache circular queue to reduce misjudgment and improve the recognition efficiency.

[0094] In some embodiments of the present application, S220 may include but is not limited to S221 to S222:

[0095] S221. When the confidence is less than the minimum threshold in the set threshold range, then eliminate the sign language isolated word corresponding to the confidence.

[0096] S222. When the confidence level is greater than the maximum threshold in the set threshold range, the sign language isolated word corresponding to the confidence level is used as a single isolated word information.

[0097] In an embodiment S221, if the confidence level is less than the minimum threshold in the set threshold range, the image information below the minimum threshold is regarded as noise data and eliminated.

[0098] In an embodiment, if the confidence level of the current first image information is less than 0.7, the current first image information is eliminated and the next set of first image information is selected.

[0099] In an embodiment S221, if the confidence level is greater than the maximum threshold in the set threshold range, the sign language isolated word inferred and recognized from the corresponding image information is recorded as isolated word information.

[0100] In an embodiment, if the confidence level of the current first image information is greater than 0.85, the sign language isolated word corresponding to the confidence level is recorded as a single isolated word information.

[0101] Through the above solution, the isolated word information is screened by the confidence level to improve the accuracy of gesture recognition and the processing speed.

[0102] In some embodiments of the present application, S400 may include but is not limited to S410 to S440:

[0103] S410. Take the read isolated word information as the first input feature and the read context information as the second input feature.

[0104] S420. According to the set semantic integration large model, obtain the set prompt word template, and determine the context information according to the first input feature and the set prompt word template.

[0105] S430. Input the context information, the set prompt word template and the second input feature into the set semantic integration large model to semantically integrate the first input feature and obtain the output feature.

[0106] S440. Analyze the output feature to obtain the used isolated word information and the output sentence containing the used isolated word information.

[0107] In an embodiment S410, call the set semantic integration large model, take the isolated word information read in time series as the first input feature, and take the read context information as the second input feature.

[0108] Among them, if there is no context information at the initial reading, the second input feature is empty.

[0109] In an embodiment S420, through the set semantic integration large model, obtain the prompt template preset through prompt engineering. Through the set prompt template and the first input feature, the set semantic integration large model performs classification evaluation to capture context information.

[0110] Among them, context information can also be accurately captured through the set prompt template, the first input feature, and the second input feature.

[0111] In an embodiment S430, in the set semantic integration large model, through the context information, combined with the set prompt template preset by prompt engineering, and related to the context information, semantic integration is performed on the isolated word sequence with time series, missing words are supplemented, and the output feature is obtained.

[0112] In an embodiment S440, since the output format is set to the json format, the output feature output by the set semantic integration large model is parsed, and the output format includes the used isolated word information and the integrated output statement.

[0113] The output feature is parsed to obtain the used isolated word information and the output statement containing the isolated word information.

[0114] Among them, the set semantic integration large model can be the Baidu Qianfan large language model or other large language models. In this application, the specific type of the set semantic integration large model is not limited.

[0115] Based on the above embodiments, exemplarily, when calling the set semantic integration large model, the large model is adjusted through prompt engineering. The prompt template "Prompt" is preset as gesture semantic analysis, and the output format of the large model is set to the json format, including the used isolated word information "input_word" and the integrated output statement "sentence". Read the isolated word information in the cache circular queue and use it as the first input feature "Isolated_words", and read the context information and use it as the second input feature "context". Combine the set prompt template "Prompt" and the first input feature "Isolated_words" to determine the context information. Based on the context information, combine the set prompt template "Prompt", the first input feature "Isolated_words", and the second input feature "context", and input them into the set semantic integration large model for processing. The set semantic integration large model performs semantic integration to obtain the corresponding result, that is, the output feature in json format, and parse the output feature to obtain the used isolated word information "input_word" and the integrated output statement "sentence".

[0116] In some embodiments of the present application, S400 may include but is not limited to S450:

[0117] S450, deleting the used isolated word information from the cache circular queue;

[0118] S460, storing the output statement in text form in the current stored text to update the context information.

[0119] In one embodiment of S450, the corresponding isolated word information in the cache circular queue is deleted through the parsed used isolated word information.

[0120] In one embodiment of S460, through the parsed output statement, it is displayed or played by voice, and saved in text form to the current stored text to update the context information.

[0121] In some embodiments of the present application, the control method further includes:

[0122] S600, obtaining second image information, and detecting whether there is vehicle information in the second image information according to the set vehicle detection model;

[0123] S610, if so, then according to the set vehicle detection model, processing the second image information, outputting the bounding box coordinates of the vehicle information, obtaining the structural parameters of the wide-angle image collector, and estimating the position information of the vehicle information according to the structural parameters and the bounding box coordinates;

[0124] S620, according to the set vehicle detection model, comparing the bounding box coordinates corresponding to the vehicle information in consecutive frames, and estimating the motion trend of the vehicle information according to the corresponding bounding box coordinates;

[0125] S630, displaying the position information and the motion trend in the waveguide display module.

[0126] In some embodiments of S600, the second image information is obtained through the wide-angle image collector, and the second image information is used as an input feature and input into the set vehicle detection model, and the set vehicle detection model is used to detect whether there is vehicle information.

[0127] In some embodiments of S610, when it is determined that there is vehicle information, then the second image information is processed through the set vehicle detection model to obtain the bounding box coordinates of the vehicle information.

[0128] Based on the installation position of the wide-angle image collector on the vision glasses, the orientation of vehicle information can be initially estimated, that is, it can be initially determined that the vehicle information comes from the rear. Based on the installation position of the wide-angle image collection, through the structural parameters and bounding box coordinates of the wide-angle image collector, the oncoming vehicle distance of the vehicle information to the user and the oncoming vehicle orientation of the vehicle information relative to the user are estimated. The position information includes: oncoming vehicle distance and oncoming vehicle orientation.

[0129] In some embodiments S620, through the set vehicle detection model, according to the bounding box coordinates corresponding to the vehicle information in consecutive frames, the motion trend of the vehicle information is estimated.

[0130] Among them, the set vehicle detection model is the PP-YOLOE vehicle detection model.

[0131] In some embodiments S630, when it is determined that there is vehicle information, a warning message is generated. According to the position information and the motion trend, the driving state of the current vehicle information is determined. Through this driving state, the warning level is determined. According to this warning level, the corresponding rendered color identifier is determined.

[0132] According to this warning message, the set display identifier area in the waveguide display module is determined, and the color identifier corresponding to this warning level is rendered and displayed in the display identifier area.

[0133] That is to say, referring to Figure 4 , through the set vehicle detection model and the second image information, it is detected whether there is an oncoming vehicle from the rear. If so, a warning message is generated. Based on the installation position of the wide-angle image collector on the vision glasses and the second image information, the position information and motion trend of the oncoming vehicle from the rear are estimated. Through the position information and the motion trend, the corresponding warning level is determined, and the color warning box corresponding to this warning level is determined. The color warning box corresponding to this position information and motion trend is rendered and displayed in the display identifier area, that is, the upper right corner area of the waveguide display module, so as to remind the deaf-mute group of the oncoming vehicle distance, orientation and motion trend from the rear, and prevent the travel difficulties caused by the hearing defect of the deaf-mute group.

[0134] Through the above solution, for the travel needs of the deaf-mute group, a multi-level warning mechanism is adopted, and a depth estimation algorithm is designed to judge information such as the distance, orientation and motion trend of oncoming vehicles from the rear, so as to provide a hierarchical warning for the user. This warning mechanism can not only effectively reduce the risk of traffic accidents encountered by the deaf-mute group during travel, but also help this group make a quick response in a complex traffic environment through intuitive visual prompts. Provide travel service assistance functions for the deaf-mute group, and significantly improve travel safety and quality of life.

[0135] In some embodiments of the present application, the control method further includes:

[0136] S710. When it is confirmed that there is a gesture action in the current first image information, the gesture action is compared with the set voice playback gesture action.

[0137] S711. When it is confirmed that the gesture action is the set voice playback gesture action, a voice playback instruction is obtained, and according to the voice playback instruction, the output statement is played.

[0138] In an embodiment of S710, when it is confirmed that there is a gesture action, the gesture action is recognized and compared with the set voice playback gesture action to determine whether to activate the voice playback module.

[0139] In an embodiment of S711, when it is determined that the gesture action is the set voice playback gesture action, a voice playback instruction is obtained and the voice playback mode is activated.

[0140] In the running voice playback mode, when the output statement is obtained through output feature parsing, the display module directly reads and displays the output statement, completes voice synthesis, and outputs and plays the output statement.

[0141] Among them, in the running voice playback mode, a warning message can also be played.

[0142] In some embodiments of the present application, the control method further includes:

[0143] S720. When it is confirmed that there is a gesture action in the current first image information, the gesture action is compared with the set voice activation gesture action.

[0144] S721. When it is confirmed that the gesture action is the set voice activation gesture action, a voice recognition activation instruction is obtained and the voice recognition module is activated.

[0145] S722. According to the voice recognition activation instruction, environmental voice information is obtained, the environmental voice information is converted into text information, and the text information is displayed in the waveguide display module or the text information is played.

[0146] In an embodiment of S720, when it is confirmed that there is a gesture action, the gesture action is recognized and compared with the set voice activation gesture action to determine whether to activate the voice recognition module.

[0147] In an embodiment of S721, when it is determined that the gesture action is the set voice playback gesture action, environmental voice information is collected through the voice recognition module, the voice information is converted into text information, and the text information is displayed or voice broadcast through the waveguide display module.

[0148] Through the above solution, by setting up the voice recognition module and the voice playback mode, voice recognition and synthesis can be achieved, the interactivity of the visual glasses is improved, two-way interaction between the user and the talker is realized, thereby improving the interactivity and user experience.

[0149] An embodiment of the present application further provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above control method for the deaf-mute travel assistance vision glasses. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0150] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented by the device embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0151] Please refer to Figure 5 , Figure 5 , which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0152] A processor 501, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0153] A memory 502, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 502 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 502, and the processor 501 is used to call and execute the control method for the deaf-mute travel assistance vision glasses in the embodiments of the present application;

[0154] An input / output interface 503, which is used to implement information input and output;

[0155] A communication interface 504, which is used to implement communication interaction between the device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.);

[0156] A bus 505, which transmits information between various components of the device (such as the processor 501, the memory 502, the input / output interface 503, and the communication interface 504);

[0157] Among them, the processor 501, the memory 502, the input / output interface 503, and the communication interface 504 are communicatively connected to each other inside the device through the bus 505.

[0158] The embodiment of the present application also provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above control method based on the deaf-mute travel assistance vision glasses.

[0159] It can be understood that the content in the above method embodiments is applicable to the present storage medium embodiment. The functions specifically implemented by the present storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0160] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0161] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0162] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown, or combine certain steps, or different steps.

[0163] Those of ordinary skill in the art can understand that all or some of the steps in the above-disclosed methods, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0164] In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules does not necessarily have to be limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0165] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more.

[0166] The modules described above as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0167] In addition, the functional modules in each embodiment of this application can be integrated into a processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.

[0168] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, and thus do not limit the scope of rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of rights of the embodiments of this application.

Claims

1. A control method for a visual aid glasses for deaf-mute people to travel, characterized in that, The visual glasses include a waveguide display module, a front image collector, a frame, and temple arms; One side of the frame is provided with a fixing groove, the front image collector is disposed in the fixing groove, the front image collector is used to obtain first image information, and the waveguide display module is disposed at the front of the frame; The control method includes: Obtain the current first image information, and according to the set gesture recognition model, detect whether there is a gesture action in the current first image information; If so, according to the set gesture recognition model, recognize the current first image information, extract single isolated word information, and add the isolated word information to the cache circular queue one by one according to the time sequence; When the number of the isolated word information in the cache circular queue is greater than the set quantity threshold, read the context information stored in the current stored text according to the current isolated word information; Process the isolated word information and the context information according to the set semantic integration large model, capture the context information, and perform semantic integration on the isolated word information according to the context information and the context information to obtain an integrated output statement; Display the output statement in text form on the waveguide display module.

2. The control method according to claim 1, wherein The recognizing the current first image information according to the set gesture recognition model and extracting single isolated word information includes: Perform recognition and inference on the current first image information through the set gesture recognition model, recognize sign language isolated words, and obtain the confidence of the sign language isolated words; When the confidence is within the set threshold range, recognize the current first image information multiple times within the set time period, obtain the corresponding confidence and record the current recognition times until the current recognition times reach the set times threshold; When the mean value of the corresponding confidence is greater than the set confidence threshold, use the sign language isolated word as the single isolated word information.

3. The control method according to claim 1, wherein The integration process of the output statement includes: Use the read isolated word information as the first input feature, and use the read context information as the second input feature; According to the set semantic integration large model, obtain the set prompt word template, and determine the context information according to the first input feature and the set prompt word template; Input the context information, the set prompt word template, and the second input feature into the set semantic integration large model to perform semantic integration on the first input feature to obtain an output feature; Analyze the output feature to obtain the used isolated word information and the output statement containing the used isolated word information.

4. The control method according to claim 3, wherein The control method further includes: Delete the used isolated word information from the cache circular queue; Store the output statement in text form in the current stored text to update the context information.

5. The control method according to claim 1, characterized in that The visual glasses further include a wide-angle image collector, the wide-angle image collector is disposed at the end of one of the temple arms, the wide-angle image collector is used to obtain second image information, and the control method further includes: Obtain the second image information, and according to the set vehicle detection model, detect whether there is vehicle information in the second image information; If so, process the second image information according to the set vehicle detection model, output the bounding box coordinates of the vehicle information, obtain the structural parameters of the wide-angle image collector, and estimate the position information of the vehicle information according to the structural parameters and the bounding box coordinates; According to the set vehicle detection model, compare the bounding box coordinates corresponding to the vehicle information in consecutive frames, and estimate the motion trend of the vehicle information according to the corresponding bounding box coordinates; Display the position information and the motion trend in the waveguide display module.

6. The control method according to claim 1, characterized in that The control method further includes: When it is confirmed that there is a gesture action in the current first image information, compare the gesture action with the set voice playback gesture action; When it is confirmed that the gesture action is the set voice playback gesture action, obtain a voice playback instruction, and play the output statement according to the voice playback instruction.

7. The control method according to claim 6, characterized in that, The visual glasses further include a voice recognition module, and the control method further includes: When it is confirmed that there is a gesture action in the current first image information, compare the gesture action with the set voice activation gesture action; When it is confirmed that the gesture action is the set voice activation gesture action, obtain a voice recognition activation instruction and activate the voice recognition module; According to the voice recognition activation instruction, obtain ambient voice information, convert the ambient voice information into text information, and display the text information in the waveguide display module or play the text information.

8. The control method according to claim 2, characterized in that The "when the confidence level is within the set threshold range" includes: When the confidence level is less than the minimum threshold in the set threshold range, eliminate the sign language isolated word corresponding to the confidence level.

9. An electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 8 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that, The computer program, when executed by the processor, implements the method according to any one of claims 1 to 8.