Multi-modal interaction aging-adaptive interface self-adaptive system

Through the multimodal interactive aging-friendly interface adaptive system, the problems of dialect identification disorders, poor tactile perception and high erroneous touch rate among elderly users are solved, and more accurate dialect recognition and tactile feedback are achieved, improving the user experience and health and safety of elderly users.

CN120335876APending Publication Date: 2025-07-18AIMI (BEIJING) ROBOT CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510438151.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing smart devices have problems such as dialect recognition disorders, poor tactile perception and high mistouch rate when used by elderly users, which affect the health and user experience of the elderly.

Method used

A multimodal interactive aging-friendly interface adaptive system is adopted, including a multimodal input module, a tactile feedback module and a dual redundant wake-up module. Data is collected through multimodal integrated sensors for dialect recognition and tactile perception, and the trigger effectiveness is judged through a false touch suppression algorithm.

Benefits of technology

It improves the accuracy of smart devices in identifying dialects of elderly users, improves tactile perception, reduces the rate of false touch, and improves the user experience and health protection of elderly users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335876A_ABST
    Figure CN120335876A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode interaction adaptive system for an aging-adaptive interface. The multi-modal interaction aging-adaptive interface self-adaption system comprises a multi-modal input module, a tactile feedback module and a dual-redundancy wake-up module. Wherein the multi-modal input module is used for acquiring current joint description data acquired by the multi-modal integrated sensor, and performing identification and dialect adaptation processing on the current joint description data to obtain a dialect identification result corresponding to a target user; the tactile feedback module is used for analyzing the current joint description data when receiving a tactile feedback instruction to obtain a tactile perception result; and the dual-redundancy wake-up module is used for judging whether the triggering of the current joint description data is valid or not through a preset false touch suppression algorithm to obtain a triggering judgment result. The dialect recognition method solves the problems that an intelligent device cannot recognize dialects more accurately, tactile perception is poor and the mistaken touch rate is high, dialect recognition is better achieved, health of old people is guaranteed, and the experience feeling of the old people using the intelligent device is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a multi-modal interactive aging-friendly interface adaptive system. Background Art

[0002] With the rapid development of science and technology, smart devices are increasingly used in people's lives. However, for the elderly, the existing smart device interaction systems have many technical defects, which brings great inconvenience to their use.

[0003] In the process of realizing the present invention, the inventors found that the prior art has the following defects: At present, when elderly users use the voice command function of smart devices, there is a problem of dialect recognition barrier, which cannot serve the elderly well. In addition, physical buttons or touch screen operations pose a great safety hazard to some disabled elderly people. Since disabled elderly people may have problems such as poor limb coordination and insufficient strength, they are prone to fall due to loss of balance when operating the device, posing a serious threat to their physical health. In addition, the wake-up mechanism of current smart devices has a high false touch rate. The high false touch rate will not only lead to unnecessary responses of the device and waste resources, but also seriously affect the user experience, causing elderly users to be troubled and dissatisfied with smart devices. Summary of the invention

[0004] The present invention provides a multimodal interactive aging-friendly interface adaptive system to achieve better dialect recognition, protect the health of the elderly, and improve the elderly's experience in using smart devices.

[0005] According to one aspect of the present invention, a multimodal interactive aging-friendly interface adaptive system is provided, wherein the multimodal interactive aging-friendly interface adaptive system comprises: a multimodal input module, a tactile feedback module and a dual redundant wake-up module;

[0006] The multimodal input module is used to obtain the current joint description data collected by the multimodal integrated sensor, and perform recognition and dialect adaptation processing on the current joint description data to obtain the dialect recognition result corresponding to the target user;

[0007] A tactile feedback module, configured to parse the current joint description data to obtain a tactile perception result when receiving a tactile feedback instruction;

[0008] The dual redundant wake-up module is used to determine whether the trigger of the current joint description data is valid through a preset false touch suppression algorithm to obtain a trigger determination result.

[0009] Further, it includes: The multimodal integrated sensor includes a camera sensor, a microphone sensor, and a piezoelectric ceramic sensor; wherein, the current joint description data includes current image data, current voice data, and current physical trigger fall action data; wherein, the camera sensor is used to collect the current image data; the microphone sensor is used to collect the current voice data; the piezoelectric ceramic sensor is used to detect the current physical trigger fall action data.

[0010] Further, the multimodal input module includes a text recognition unit, a speech recognition unit, and a dialect adaptation unit; wherein, the text recognition unit is used to perform text recognition on the obtained current image data through a pre-set convolutional recurrent neural network to obtain a text recognition result; the speech recognition unit is used to perform speech recognition on the obtained current voice data through a pre-set Wav2Vec2 model to obtain a speech recognition result; the dialect adaptation unit is used to receive and perform dialect adaptation processing based on the text recognition result and the speech recognition result to obtain a dialect recognition result corresponding to the target user.

[0011] Further, it includes: The dialect adaptation unit is further used to perform fusion processing on the received text recognition result and the speech recognition result, and perform prediction on the obtained fusion processing result through a pre-set dialect adaptation model to obtain the dialect recognition result.

[0012] Further, the tactile feedback module includes: an ultrasonic signal generation unit and a tactile perception result generation unit; wherein, the ultrasonic signal generation unit is used to generate a frequency-pressure ultrasonic signal according to the current physical trigger fall action data when receiving a tactile feedback instruction; the tactile perception result generation unit is used to analyze the frequency-pressure ultrasonic signal to obtain a tactile perception result.

[0013] Further, the tactile perception result generation unit includes an environmental noise cancellation subunit and a tactile perception result generation subunit; the multimodal integrated sensor is further used to collect current environmental noise; wherein, the environmental noise cancellation subunit is used to perform real-time monitoring and analysis on the obtained current environmental noise, and perform environmental noise cancellation processing on the frequency-pressure ultrasonic signal according to the adjusted filtering parameters to obtain a target frequency-pressure ultrasonic signal; the tactile perception result generation subunit is used to obtain a perceivable tactile feedback pattern in the target area and generate a tactile perception result according to the target frequency-pressure ultrasonic signal.

[0014] Further, the haptic feedback module further includes: a behavior habit storage unit; wherein the behavior habit storage unit is configured to obtain the identity number of the target user, and perform a joint storage operation on the obtained current physical trigger fall action data and the haptic perception result, and store them in the behavior habit database corresponding to the behavior habit storage unit.

[0015] Further, the current joint description data further includes: a current timestamp, a current acceleration, and a current speech confidence level.

[0016] Further, the dual-redundancy wake-up module includes: a data receiving unit, a trigger validity judgment unit, and a trigger validity result feedback unit; wherein, the data receiving unit is configured to receive the current timestamp, the current acceleration, and the current speech confidence level; the trigger validity judgment unit is configured to respectively and sequentially judge whether the received current timestamp, the current acceleration, and the current speech confidence level are valid triggers through a pre-set false touch suppression algorithm, and obtain a trigger determination result; the trigger validity result feedback unit is configured to, if it is determined that the trigger determination result is a trigger valid result, perform a feedback process on the trigger valid result.

[0017] Further, it includes: the trigger validity judgment unit is further configured to: obtain a timestamp threshold, an acceleration threshold, and a speech confidence level threshold, and obtain the previous trigger timestamp corresponding to the current timestamp; calculate the difference between the current timestamp and the previous trigger timestamp, and compare it with the timestamp threshold. If the difference is less than the timestamp threshold, continue to judge whether the norm of the current acceleration is less than the acceleration threshold. If so, continue to judge whether the current speech confidence level is less than the speech confidence level threshold. If so, determine that the trigger determination result is a trigger invalid result; if any one of the difference is not less than the timestamp threshold, the current acceleration is not less than the acceleration threshold, or the current speech confidence level is not less than the speech confidence level threshold is satisfied, determine that the trigger determination result is a trigger valid result.

[0018] In the technical solution of the embodiment of the present invention, the multi-modal interaction aging-friendly interface adaptive system includes a multi-modal input module, a haptic feedback module, and a dual-redundancy wake-up module; among them, the multi-modal input module is used to obtain the current joint description data collected by the multi-modal integrated sensor, and perform identification and dialect adaptation processing on the current joint description data to obtain the dialect recognition result corresponding to the target user; the haptic feedback module is used to parse the current joint description data to obtain a haptic perception result when receiving a haptic feedback instruction; the dual-redundancy wake-up module is used to judge whether the trigger is effective for the current joint description data through a pre-set accidental touch suppression algorithm to obtain a trigger determination result. It solves the problems that intelligent devices cannot more accurately recognize dialects, have poor haptic perception, and a high accidental touch rate, better realizes the recognition of dialects, protects the health of the elderly, and improves the experience of the elderly using intelligent devices.

[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Brief Description of the Drawings

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 It is a schematic structural diagram of a multi-modal interaction aging-friendly interface adaptive system provided in Embodiment 1 of the present invention;

[0022] Figure 2 It is a detailed schematic structural diagram of the multi-modal input module in a multi-modal interaction aging-friendly interface adaptive system provided in Embodiment 2 of the present invention;

[0023] Figure 3 It is a detailed schematic structural diagram of the haptic feedback module in a multi-modal interaction aging-friendly interface adaptive system provided in Embodiment 3 of the present invention;

[0024] Figure 4 It is a detailed schematic structural diagram of the dual-redundancy wake-up module in a multi-modal interaction aging-friendly interface adaptive system provided in Embodiment 4 of the present invention. Detailed Embodiments

[0025] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0026] It should be noted that the terms "target", "current", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0027] Embodiment 1

[0028] Figure 1 FIG. 10 is a schematic structural diagram of a multimodal interaction aging-friendly interface adaptive system provided in Embodiment 1 of the present invention. The multimodal interaction aging-friendly interface adaptive system 100 includes: a multimodal input module 110, a tactile feedback module 120, and a dual-redundancy wake-up module 130.

[0029] Among them, the multimodal input module 110 is used to obtain the current joint description data collected by the multimodal integrated sensor, and perform identification and dialect adaptation processing on the current joint description data to obtain the dialect recognition result corresponding to the target user.

[0030] Among them, the multimodal integrated sensor is used to collect various types of data. Specifically, the multimodal integrated sensor may include a camera sensor, a microphone sensor, and a piezoelectric ceramic sensor.

[0031] Among them, the current joint description data includes current image data, current voice data, and current physical trigger fall action data; among them, the camera sensor is used to collect current image data; the microphone sensor is used to collect current voice data; the piezoelectric ceramic sensor is used to detect current physical trigger fall action data.

[0032] Specifically, the multimodal input module may be a module that processes various received data to implement the recognition of the dialects of different target users.

[0033] Among them, the camera sensor is equipped with a dual-camera module (a 5-megapixel main camera for image acquisition, and also equipped with a spectral confocal displacement sensor for obtaining depth information), which can clearly capture images containing text and obtain depth data of the relevant environment, providing support for multi-modal data acquisition. The piezoelectric ceramic sensor has a waterproof function and can work stably in various environments, and can detect the physical trigger operations of users. Such as when an emergency button is pressed, and it can also detect abnormal actions such as falling.

[0034] The tactile feedback module 120 is used to parse the current combined description data to obtain a tactile perception result when receiving a tactile feedback instruction.

[0035] Among them, the tactile feedback module can parse the received data to obtain the tactile perception results corresponding to different target users.

[0036] The dual-redundancy wake-up module 130 is used to judge whether the trigger of the current combined description data is effective through a preset false touch suppression algorithm, and obtain a trigger determination result.

[0037] Among them, the dual-redundancy wake-up module can be a module for judging whether the target user is effectively triggered.

[0038] In this embodiment, the problems of dialect recognition of the target user, tactile perception, and effective trigger determination can be solved in sequence through the multi-modal input module, the tactile feedback module, and the dual-redundancy wake-up module, so that the intelligent device can be better suitable for the use of the elderly, improve the experience of the elderly, and ensure the health of the elderly.

[0039] In the technical solution of the embodiment of the present invention, the multi-modal interaction aging-friendly interface adaptive system includes a multi-modal input module, a tactile feedback module, and a dual-redundancy wake-up module; among them, the multi-modal input module is used to obtain the current combined description data collected by the multi-modal integrated sensor, and perform recognition and dialect adaptation processing on the current combined description data to obtain the dialect recognition result corresponding to the target user; the tactile feedback module is used to parse the current combined description data to obtain a tactile perception result when receiving a tactile feedback instruction; the dual-redundancy wake-up module is used to judge whether the trigger of the current combined description data is effective through a preset false touch suppression algorithm, and obtain a trigger determination result. It solves the problems that intelligent devices cannot more accurately recognize dialects, have poor tactile perception, and have a high false touch rate, better realizes the recognition of dialects, ensures the health of the elderly, and improves the experience of the elderly using intelligent devices.

[0040] Embodiment 2

[0041] Figure 2This is a detailed structural schematic diagram of the multimodal input module in a multimodal interaction aging-friendly interface adaptive system provided in the second embodiment of the present invention. The multimodal input module 110 includes: a text recognition unit 111, a speech recognition unit 112, and a dialect adaptation unit 113.

[0042] Among them, the text recognition unit 111 is used to perform text recognition on the obtained current image data through a pre-set convolutional recurrent neural network to obtain a text recognition result;

[0043] The speech recognition unit 112 is used to perform speech recognition on the obtained current speech data through a pre-set Wav2Vec2 model to obtain a speech recognition result;

[0044] The dialect adaptation unit 113 is used to receive and perform dialect adaptation processing based on the text recognition result and the speech recognition result to obtain a dialect recognition result corresponding to the target user.

[0045] Among them, the text recognition unit can be a unit that performs text recognition on the collected image data. The speech recognition unit can be a unit that performs speech recognition on the collected speech data. The dialect adaptation unit can be a unit that matches the parsed data with a dialect adaptation model to obtain dialect recognition.

[0046] In this embodiment, for the text recognition unit, through the convolutional recurrent neural network in OCR (Optical Character Recognition), text recognition is performed on the current image data to obtain a text recognition result.

[0047] Among them, OCR is a technology that converts text in an image into editable and searchable text. It analyzes the character shapes in an image or scanned document, recognizes the corresponding text content, and outputs it in a computer-processable text format (such as text formats like TXT, PDF, and Word).

[0048] The advantage of such a setting is that through the convolutional recurrent neural network, efficient feature extraction and recognition of printed text can be performed, with strong adaptability and accuracy.

[0049] In this embodiment, the speech recognition unit performs speech recognition on the obtained current speech data through a pre-set Wav2Vec2 model and outputs the obtained speech recognition result in text form.

[0050] Optionally, it includes: the dialect adaptation unit is further configured to perform fusion processing on the received text recognition result and the speech recognition result, and predict the obtained fusion processing result through a pre-set dialect adaptation model to obtain the dialect recognition result.

[0051] Among them, the dialect adaptation model can be trained based on a large amount of dialect data and can perform dialect recognition and adaptation on the fused text and speech information.

[0052] Specifically, the dialect adaptation model can recognize eight major dialects such as Cantonese, Sichuan dialect, Minnan dialect, Hakka dialect, Wu dialect, Hunan dialect, Gan dialect, and Northern dialect. This can enable elderly users to obtain more accurate responses when interacting with smart devices in dialect, greatly improving the convenience of interaction.

[0053] In this embodiment, through in-depth mining and analysis of a large amount of printed text (i.e., the recognized text recognition result) and corresponding dialect speech data (i.e., the recognized speech recognition result), a fusion processing result is obtained, and a semantic association fusion model can be constructed according to the fusion processing result. This model can understand the semantic expression of printed text in different dialect speeches, realize the accurate matching of text and speech, thereby improving the accuracy of dialect recognition.

[0054] In addition, it also supports dynamic updating of the dialect speech library. When there are new dialect data or speech samples, the system can automatically incorporate them into the training and learning scope, which can be used to retrain and optimize the dialect adaptation model. This enables the system to continuously adapt to new dialect changes and user needs and maintain a high dialect recognition performance.

[0055] In the technical solution of the embodiment of the present invention, the multimodal input module includes a text recognition unit, a speech recognition unit, and a dialect adaptation unit. Through the combined processing operations of the text recognition unit, the speech recognition unit, and the dialect adaptation unit, it can better realize the recognition of dialects, improve the accuracy of dialect recognition, enable the system to continuously adapt to new dialect changes and user needs, maintain a high dialect recognition performance, and improve the experience of elderly users using smart devices.

[0056] Embodiment III

[0057] Figure 3 It is a detailed structural schematic diagram of a tactile feedback module in a multimodal interaction aging-friendly interface adaptive system provided by Embodiment III of the present invention. The tactile feedback module 120 includes: an ultrasonic signal generation unit 121 and a tactile perception result generation unit 122.

[0058] Among them, the ultrasonic signal generation unit 121 is configured to generate a frequency-pressure ultrasonic signal according to the current physical trigger fall action data when receiving a tactile feedback instruction.

[0059] The tactile perception result generation unit 122 is configured to analyze the frequency-pressure ultrasonic signal to obtain a tactile perception result.

[0060] Optionally, the tactile perception result generation unit includes an environmental noise cancellation subunit and a tactile perception result generation subunit; the multi-modal integrated sensor is further configured to collect the current environmental noise; wherein, the environmental noise cancellation subunit is configured to perform real-time monitoring and analysis on the obtained current environmental noise, and perform environmental noise cancellation processing on the frequency-pressure ultrasonic signal according to the adjusted filtering parameters to obtain a target frequency-pressure ultrasonic signal; the tactile perception result generation subunit is configured to obtain a perceivable tactile feedback pattern in the target area and generate a tactile perception result according to the target frequency-pressure ultrasonic signal.

[0061] Among them, the current physical trigger fall action data may be related data describing the user's physical trigger fall action.

[0062] In this embodiment, the tactile feedback module is a non-contact tactile interaction system based on ultrasonic technology. The system includes an ultrasonic array, which can receive a control instruction to generate an ultrasonic signal with a corresponding frequency and pressure, and can eliminate environmental noise interference through adaptive filtering, form a perceivable tactile feedback pattern in the target area, that is, obtain a tactile perception result, and can be programmed to generate multiple tactile modes, supporting more than 256 different interaction modes. Among them, the ultrasonic array adopts a 16×16 contact density, which can precisely control the emission and propagation of ultrasonic waves and form a high-precision tactile feedback pattern in the target area.

[0063] Furthermore, a programmable tactile coding library can be set. Developing this programmable tactile coding library can support more than 256 interaction modes. This coding library allows developers to flexibly program various different tactile feedback modes according to different application scenarios and user requirements, such as simulating button clicks, swipes, and reminders and other interaction actions, greatly enriching the application scenarios and interaction methods of tactile feedback.

[0064] Correspondingly, for the non-contact tactile interaction system, it realizes non-contact tactile control with centimeter-level precision and meets the vibration standard of medical devices. By precisely controlling the emission and propagation of ultrasonic signals, a highly accurate tactile feedback pattern can be formed in the target area, and its precision can be accurate to the centimeter level. This precision not only provides more accurate operation feedback for users, but also meets the strict standards of medical devices in terms of tactile perception, ensuring the safety and reliability of tactile feedback.

[0065] Among them, the environmental noise cancellation subunit can be a subunit that cancels environmental noise through adjusted filtering parameters. The tactile perception result generation subunit can be a unit that generates specific user tactile perceptions.

[0066] In this embodiment, for the specific tactile perception result feedback, the tactile feedback frequency range can be set between 20 - 200 Hz. This frequency range can generate tactile sensations of different intensities and textures, and can simulate both gentle touches and relatively strong vibrations. The pressure intensity can be set between 0.1 - 0.5 bar to ensure obvious tactile feedback for the user within a safe range. The acting area is 5 - 15 cm 2 , which can generate perceivable tactile sensations in specific areas of the user's body without causing discomfort to the user.

[0067] Optionally, the tactile feedback module further includes: a behavior habit storage unit; among them, the behavior habit storage unit is used to obtain the identity number of the target user, and perform a joint storage operation on the obtained current physical trigger fall action data and the tactile perception result, and store them in the behavior habit database corresponding to the behavior habit storage unit.

[0068] In this embodiment, different behavior habits can be constructed according to different target users, and storage or re - learning operations can be performed according to different behavior habits, so as to achieve more accurate tactile feedback.

[0069] The technical solution of the embodiment of the present invention, through the specific setting of the tactile feedback module, specifically, the tactile feedback module can include an ultrasonic signal generation unit and a tactile perception result generation unit. In this way, the feedback of tactile perception results can be more accurate, which can further ensure the health of the elderly. The construction of the behavior habit database can perform more accurate tactile feedback according to the user's historical behavior habits, ensuring the safety and reliability of tactile feedback.

[0070] Embodiment Four

[0071] Figure 4 This is a detailed structural schematic diagram of a dual - redundant wake - up module in a multi - modal interaction aging - friendly interface adaptive system provided by Embodiment Four of the present invention. The dual - redundant wake - up module 130 includes: a data receiving unit 131, a trigger validity judgment unit 132, and a trigger valid result feedback unit 133.

[0072] Among them, the current combined description data further includes: a current timestamp, a current acceleration, and a current speech confidence level.

[0073] Optionally, the data receiving unit is configured to receive the current timestamp, the current acceleration, and the current speech confidence level; the trigger validity judgment unit is configured to respectively and sequentially judge the validity of the trigger for the received current timestamp, the current acceleration, and the current speech confidence level through a preset accidental touch suppression algorithm to obtain a trigger determination result; the trigger valid result feedback unit is configured to, if it is determined that the trigger determination result is a trigger valid result, perform feedback processing on the trigger valid result.

[0074] Optionally, it includes: the trigger validity judgment unit is further configured to: obtain a timestamp threshold, an acceleration threshold, and a speech confidence level threshold, and obtain the previous trigger timestamp corresponding to the current timestamp; calculate the difference between the current timestamp and the previous trigger timestamp, and compare it with the timestamp threshold. If the difference is less than the timestamp threshold, continue to judge whether the norm of the current acceleration is less than the acceleration threshold. If so, continue to judge whether the current speech confidence level is less than the speech confidence level threshold. If so, determine that the trigger determination result is a trigger invalid result; if any one of the conditions that the difference is not less than the timestamp threshold, the current acceleration is not less than the acceleration threshold, or the current speech confidence level is not less than the speech confidence level threshold is satisfied, determine that the trigger determination result is a trigger valid result.

[0075] Among them, the accidental touch suppression algorithm determines the validity of the trigger by comprehensively judging multiple parameters such as the current timestamp, the current acceleration, and the current speech confidence level.

[0076] Exemplarily, it is assumed that the timestamp threshold is set to 2 seconds, the acceleration threshold is set to 0.5, and the speech confidence level threshold is set to 0.8.

[0077] It is assumed that the difference between the current timestamp and the previous trigger timestamp is 1 second, the norm of the current acceleration is 0.3, and the speech confidence level is 0.5.

[0078] Specifically, since the difference between the current timestamp and the previous trigger timestamp (1 second) is less than the timestamp threshold (2 seconds), it is necessary to continue to compare the norm of the current acceleration and the acceleration threshold. Since the norm of the current acceleration (0.3) is less than the acceleration threshold (0.5), continue to judge the size of the current speech confidence level and the speech confidence level threshold. Since the current speech confidence level (0.5) is less than the speech confidence level threshold (0.8), it can be determined that the trigger determination result is a trigger invalid result.

[0079] It is assumed that the difference between the current timestamp and the previous trigger timestamp is 3 seconds. Since the difference between the current timestamp and the previous trigger timestamp (3 seconds) is greater than the timestamp threshold (2 seconds), it can be determined that the trigger determination result is a trigger valid result.

[0080] Assume that the norm of the current acceleration is 0.6. Since the norm of the current acceleration (0.6) is greater than the acceleration threshold (0.5), it can be determined that the trigger determination result is a valid trigger result.

[0081] Assume that the speech confidence is 0.9. Since the current speech confidence (0.9) is greater than the speech confidence threshold (0.8), it can be determined that the trigger determination result is a valid trigger result.

[0082] The technical solution of the embodiment of the present invention reduces the unnecessary responses of the device by sequentially comparing the current timestamp, the current acceleration, and the current speech confidence, improves the response speed of the device, significantly optimizes the usage experience of elderly users, and enables them to use smart devices more easily and pleasantly.

[0083] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A multi-modal interaction age-friendly interface adaptive system, characterized in that, The multi-modal interaction aging-friendly interface adaptive system includes: a multi-modal input module, a tactile feedback module, and a dual-redundancy wake-up module; Among them, the multi-modal input module is used to obtain the current joint description data collected by the multi-modal integrated sensor, and perform recognition and dialect adaptation processing on the current joint description data to obtain the dialect recognition result corresponding to the target user; The tactile feedback module is used to parse the current joint description data to obtain a tactile perception result when receiving a tactile feedback instruction; The dual-redundancy wake-up module is used to judge whether the trigger of the current joint description data is effective through a preset accidental touch suppression algorithm to obtain a trigger determination result.

2. The system according to claim 1, wherein It includes: The multi-modal integrated sensor includes a camera sensor, a microphone sensor, and a piezoelectric ceramic sensor; Among them, the current joint description data includes current image data, current voice data, and current physical trigger fall action data; Among them, the camera sensor is used to collect current image data; the microphone sensor is used to collect current voice data; the piezoelectric ceramic sensor is used to detect current physical trigger fall action data.

3. The system according to claim 2, wherein The multi-modal input module includes a text recognition unit, a voice recognition unit, and a dialect adaptation unit; Among them, the text recognition unit is used to perform text recognition on the obtained current image data through a preset convolutional recurrent neural network to obtain a text recognition result; The voice recognition unit is used to perform voice recognition on the obtained current voice data through a preset Wav2Vec2 model to obtain a voice recognition result; The dialect adaptation unit is used to receive and perform dialect adaptation processing according to the text recognition result and the voice recognition result to obtain the dialect recognition result corresponding to the target user.

4. The system according to claim 3, wherein It includes: The dialect adaptation unit is further used to fuse the received text recognition result and the voice recognition result, and predict the obtained fusion processing result through a preset dialect adaptation model to obtain the dialect recognition result.

5. The system according to claim 2, wherein The tactile feedback module includes: an ultrasonic signal generation unit and a tactile perception result generation unit; Among them, the ultrasonic signal generation unit is used to generate a frequency-pressure ultrasonic signal according to the current physical trigger fall action data when receiving a tactile feedback instruction; The tactile perception result generation unit is used to parse the frequency-pressure ultrasonic signal to obtain a tactile perception result.

6. The system according to claim 5, wherein The tactile perception result generation unit includes an environmental noise cancellation subunit and a tactile perception result generation subunit; the multi-modal integrated sensor is further used to collect current environmental noise; Among them, the environmental noise cancellation subunit is used to perform real-time monitoring and analysis on the obtained current environmental noise, and perform environmental noise cancellation processing on the frequency-pressure ultrasonic signal according to the adjusted filtering parameters to obtain a target frequency-pressure ultrasonic signal; The tactile perception result generation subunit is used to obtain the perceivable tactile feedback pattern in the target area, and generate a tactile perception result according to the target frequency-pressure ultrasonic signal.

7. The system according to claim 6, wherein The tactile feedback module further includes: a behavior habit storage unit; Among them, the behavior habit storage unit is used to obtain the identity number of the target user, and perform a joint storage operation on the obtained current physically-triggered fall action data and the tactile perception result, and store them in the behavior habit database corresponding to the behavior habit storage unit.

8. The system according to claim 1, wherein The current joint description data further includes: a current timestamp, a current acceleration, and a current voice confidence level.

9. The system according to claim 8, wherein The dual-redundancy wake-up module includes: a data receiving unit, a trigger validity judgment unit, and a trigger validity result feedback unit; Among them, the data receiving unit is used to receive the current timestamp, the current acceleration, and the current voice confidence level; The trigger validity judgment unit is used to judge whether the received current timestamp, current acceleration, and current voice confidence level are valid triggers respectively in sequence through a pre-set false trigger suppression algorithm, and obtain a trigger determination result; The trigger validity result feedback unit is used to, if it is determined that the trigger determination result is a trigger validity result, perform feedback processing on the trigger validity result.

10. The system according to claim 9, wherein Including: The trigger validity judgment unit is further used to: Obtain a timestamp threshold, an acceleration threshold, and a voice confidence level threshold, and obtain the previous trigger timestamp corresponding to the current timestamp; Calculate the difference between the current timestamp and the previous trigger timestamp, and compare it with the timestamp threshold. If the difference is less than the timestamp threshold, continue to judge whether the norm of the current acceleration is less than the acceleration threshold. If so, continue to judge whether the current voice confidence level is less than the voice confidence level threshold. If so, determine that the trigger determination result is a trigger invalid result; If any one of the conditions that the difference is not less than the timestamp threshold, the current acceleration is not less than the acceleration threshold, or the current voice confidence level is not less than the voice confidence level threshold is satisfied, determine that the trigger determination result is a trigger validity result.

Citation Information

Cited By

  • Ageing-suitable multimedia data generation method and system based on medical treatment and transboundary fusion

    CN121601192A