Multi-modal data input method and apparatus, terminal device, and storage medium
Patent Information
- Application Number
- CA3317564
- Authority / Receiving Office
- CA · CA
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-06
- Filing Date
- 2024-10-11
- Publication Date
- 2026-09-21
AI Technical Summary
Existing smart health and elderly care products require manual input of physiological indicators for the elderly, which is cumbersome and prone to numerical input errors, resulting in a poor user experience and hindering their widespread adoption.
By identifying images or audio recordings related to physiological indicator data, image detection models and voice detection models are used to quickly extract physiological indicator data and save it to a health management database.
It enables rapid acquisition of physiological indicator data, avoids the drawbacks of manual input, and improves the ease of use of smart health and elderly care products.
Abstract
Description
A multimodal data input method, apparatus, terminal device, and storage medium Technical Field
[0001] This invention relates to the fields of image processing and speech processing technology, and in particular to a multimodal data input method, apparatus, terminal device and storage medium. Background Technology
[0002] With the current trend of population structure transformation in my country, more and more smart products will enter the home-based elderly care scenario. How to provide companionship and corresponding smart services for the elderly at home is a topic of common concern for society and enterprises. Smart health and elderly care products generally include smart detection devices, electrocardiogram monitoring devices, and blood pressure monitoring devices. These products can meet the urgent needs and health monitoring of the elderly population in a short period of time. With the development of next-generation artificial intelligence, new smart health and elderly care products can rely on technologies such as next-generation artificial intelligence models to conduct multi-dimensional analysis of various physiological indicators of the elderly through data mining, enabling potential disease risk assessment and early intervention.
[0003] However, most new smart health and elderly care products require manual input to obtain the physiological indicators of the elderly. The data for these physiological indicators is extensive, the operation is cumbersome, and manual input often results in numerical errors, leading to a very poor user experience for smart health and elderly care products and ultimately hindering their widespread adoption.
[0004] Therefore, how to enable quick data input for smart health and elderly care products has become an urgent problem to be solved.
[0005] Summary of the Invention
[0006] This invention provides a multimodal data input method, apparatus, terminal device, and storage medium. By recognizing images or audio recordings related to physiological indicator data, various physiological indicator data can be quickly acquired and saved to a health management database, avoiding the drawbacks of manual input and improving the ease of use of smart health and elderly care products.
[0007] An embodiment of the present invention provides a multimodal data input method, comprising:
[0008] Get the data input pattern;
[0009] When the data input mode is determined to be image input mode, physiological indicator data images are acquired, and data is extracted from the physiological indicator data images through a preset image detection model, and then various physiological indicators and the physiological data corresponding to each physiological indicator are output.
[0010] When the data input mode is determined to be voice input mode, the recording data is acquired, and the recording data is extracted through a preset voice detection model, and then various physiological indicators and the physiological data corresponding to each physiological indicator are output.
[0011] All physiological indicators and their corresponding physiological data are saved to a pre-set health management database.
[0012] Furthermore, prior to acquiring the data input pattern, the following is also included:
[0013] Obtain several physiological indicator fields and construct a health management data table to be populated.
[0014] Furthermore, the step of extracting data from the physiological indicator data image using a preset image detection model, and then outputting various physiological indicators and their corresponding physiological data, includes:
[0015] Right-angle edge detection is performed on the physiological indicator data image to obtain the corner points of the physiological indicator data image;
[0016] Based on the corner points, the physiological indicator data image is segmented to generate several segmented images;
[0017] The segmented image is subjected to perspective transformation to generate several images to be detected;
[0018] The image to be detected is input into a preset image detection model so that the image detection model can identify and output various physiological indicators and the physiological data corresponding to each physiological indicator.
[0019] Furthermore, the step of performing right-angle edge detection on the physiological indicator data image to obtain the corner points of the physiological indicator data image includes:
[0020] Edge detection is performed on the physiological indicator data image to obtain several edge pixels of the physiological indicator data image;
[0021] Line fitting is performed on a number of the aforementioned edge pixels to generate a number of line clusters in the physiological index data image;
[0022] The coordinates of the aforementioned line clusters are transformed to generate the polar coordinates of each line cluster;
[0023] Based on the polar coordinates of each line cluster, the edge lines of the physiological indicator data image are determined, and based on the edge lines, the corner points of the physiological indicator data image are determined.
[0024] Furthermore, the process of extracting data from the recording data using a preset speech detection model, and then outputting various physiological indicators and their corresponding physiological data, includes:
[0025] The audio data is denoised to generate audio data to be identified;
[0026] The recording data to be identified is input into the speech detection model so that the speech detection model can extract and output various physiological indicators and the corresponding physiological data from the recording data to be identified.
[0027] Furthermore, before saving the various physiological indicators and their corresponding physiological data to the preset health management database, the following steps are also included:
[0028] The system detects whether there are non-numeric symbols in each physiological data point, and if any non-numeric symbol is confirmed to exist in any physiological data point, the data is extracted again.
[0029] The system detects whether there is a decimal point in each physiological data point, and if a decimal point is confirmed in any physiological data point, the physiological data is processed into an integer.
[0030] The system detects whether there are numerical symbols in each physiological data point, and deletes the numerical symbols when they are confirmed to exist in any physiological data point.
[0031] Furthermore, the step of saving various physiological indicators and their corresponding physiological data into a preset health management database includes:
[0032] The various physiological indicators are matched with the physiological indicator fields in the health management data table, and the successfully matched physiological data is filled into the corresponding physiological indicator fields to generate a complete health management data table, which is then saved to the health management database.
[0033] Another embodiment of the present invention provides a multimodal data input device, comprising:
[0034] The mode selection module is used to acquire the data input mode;
[0035] The image detection module is used to acquire physiological indicator data images when the data input mode is determined to be image input mode, and to extract data from the physiological indicator data images through a preset image detection model, and then output various physiological indicators and the physiological data corresponding to each physiological indicator.
[0036] The voice detection module is used to acquire recording data when the data input mode is determined to be voice input mode, and to extract data from the recording data through a preset voice detection model, and then output various physiological indicators and the physiological data corresponding to each physiological indicator.
[0037] The data storage module is used to save various physiological indicators and their corresponding physiological data to a preset health management database.
[0038] Another embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement a multimodal data input method as described in any of the embodiments.
[0039] Another embodiment of the present invention provides a storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the storage medium is located to perform a multimodal data input method as described in any of the above embodiments.
[0040] The following benefits can be obtained by implementing the present invention:
[0041] This invention discloses a multimodal data input method, apparatus, terminal device, and storage medium. The method involves acquiring a data input mode; when the data input mode is determined to be an image input mode, acquiring physiological indicator data images, extracting data from the physiological indicator data images using a preset image detection model, and then outputting various physiological indicators and their corresponding physiological data, which are then saved to a preset health management database; when the data input mode is determined to be a voice input mode, acquiring recording data, extracting data from the recording data using a preset voice detection model, and then outputting various physiological indicators and their corresponding physiological data, which are then saved to the preset health management database. Therefore, this invention can quickly acquire various physiological indicator data and save them to a health management database by recognizing images or recordings related to physiological indicator data, avoiding the drawbacks of manual input and improving the ease of use of smart health and elderly care products. Attached Figure Description
[0042] Figure 1 is a flowchart illustrating a multimodal data input method according to an embodiment of the present invention.
[0043] Figure 2 is a schematic diagram of a multimodal data input device provided in an embodiment of the present invention.
[0044] Figure 3 is a flowchart of a method for image recognition input provided in an embodiment of the present invention.
[0045] Figure 4 is a flowchart of a method for using voice recognition input according to an embodiment of the present invention. Detailed Implementation
[0046] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0047] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.
[0048] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.
[0049] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0050] In the description of the embodiments in this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0051] In the description of the embodiments of this application, the term "multiple" refers to two or more (including two), similarly, "multiple sets" refers to two or more (including two sets), and "multiple pieces" refers to two or more (including two pieces).
[0052] In the description of the embodiments of this application, unless otherwise expressly specified and limited, technical terms such as "installation," "connection," "joining," and "fixing" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. For those skilled in the art, the specific meaning of the above terms in the embodiments of this application can be understood according to the specific circumstances.
[0053] Referring to Figure 1, it is a flowchart illustrating a multimodal data input method according to an embodiment of the present invention, including:
[0054] S1, Obtain data input mode;
[0055] In a preferred embodiment of the present invention, the data input mode is obtained based on the user's operation on the system's interactive interface. It should be noted that, in this embodiment, the data input mode includes: image input mode, voice input mode, and manual input mode.
[0056] Preferably, before acquiring the data input pattern, the following steps are also included:
[0057] S0. Obtain several physiological indicator fields and construct a health management data table to be populated.
[0058] In a preferred embodiment of the present invention, several physiological indicator fields are obtained based on the user's operations on the system's interactive interface, and a health management data table to be filled is constructed. It is understood that the user can select several physiological indicator fields to be entered on the interactive interface to construct the health management data table.
[0059] S2. When the data input mode is determined to be the image input mode, the physiological indicator data image is acquired, and the physiological indicator data image is extracted by a preset image detection model, and then the various physiological indicators and the physiological data corresponding to each physiological indicator are output.
[0060] In a preferred embodiment of the present invention, the physiological indicator data is captured by activating the camera device on the system, thereby obtaining images of the physiological indicator data. Specifically, the user needs to place the medical examination report or printed physiological indicator data flat within the field of view of the robot's camera, keeping it still and unobstructed, and ensuring that the ambient light is sufficient for the robot's camera to clearly capture the text in the target area. Alternatively, the user can also directly upload the captured images of the physiological indicator data.
[0061] Preferably, the step of extracting data from the physiological indicator data image using a preset image detection model, and then outputting various physiological indicators and their corresponding physiological data, includes:
[0062] S21. Perform right-angle edge detection on the physiological indicator data image to obtain the corner points of the physiological indicator data image;
[0063] Preferably, the step of performing right-angle edge detection on the physiological indicator data image to obtain the corner points of the physiological indicator data image includes:
[0064] S211. Perform edge detection on the physiological indicator data image to obtain a number of edge pixels of the physiological indicator data image;
[0065] S212. Perform line fitting on a number of the edge pixels to generate a number of line clusters in the physiological index data image;
[0066] S213. Perform coordinate transformation on the several line clusters to generate polar coordinates for each line cluster;
[0067] S214. Determine the edge lines of the physiological index data image based on the polar coordinates of each line cluster, and determine the corner points of the physiological index data image based on the edge lines.
[0068] S22. Based on the corner points, perform image segmentation on the physiological indicator data image to generate several segmented images;
[0069] S23. Perform perspective transformation on the segmented image to generate several images to be detected;
[0070] S24. Input the image to be detected into a preset image detection model so that the image detection model can identify and output various physiological indicators and the physiological data corresponding to each physiological indicator.
[0071] In a preferred embodiment of the present invention, edge detection is first performed on the physiological indicator data image to obtain all edge points of the image. The line cluster Y = kX + b of the image edge points is then converted into polar coordinate space ρ = xCosθ + ySinθ. It can be understood that for all points on any straight line in the image, there is a relatively strong signal in the polar coordinate space (ρ, θ), thereby calculating the pixel coordinates of each point on the straight line and obtaining the four edge lines of the target region. By combining the four linear equations of the four lines in two variables of the image, the four intersecting points C1, C2, C3, and C4 are solved.
[0072] Furthermore, the quadrilateral region enclosed by the four points is segmented into an image, and then transformed into a rectangle using perspective transformation. The perspective transformation formula is:
[0073] Furthermore, OCR detection and recognition are performed on the image to be detected. Specifically, in this embodiment, the lightweight deep network model PP-OCRv2 is used to perform text detection and recognition in the image, obtaining various physiological indicators for recognition and the corresponding physiological data for each indicator.
[0074] S3. When the data input mode is determined to be voice input mode, the recording data is acquired, and the recording data is extracted through a preset voice detection model, and then various physiological indicators and the physiological data corresponding to each physiological indicator are output.
[0075] In a preferred embodiment of the present invention, audio recording data is acquired by automatically monitoring the sounds broadcast by the user or detection device in the environment through the built-in microphone array. Specifically, the user selects a relatively quiet environment, turns on the system to activate the microphone array to monitor the surrounding speech, i.e., the sound of the user reading a physical examination report or physiological indicator test results, or the sound of the human indicator detection device broadcasting measurement results. Furthermore, the user can also directly upload pre-recorded audio data.
[0076] Preferably, the step of extracting data from the recording data using a preset speech detection model, and then outputting various physiological indicators and the corresponding physiological data for each physiological indicator, includes:
[0077] S31. Noise reduction is applied to the recorded data to generate recorded data to be identified;
[0078] S32. Input the recording data to be identified into the speech detection model so that the speech detection model extracts and outputs various physiological indicators and the physiological data corresponding to each physiological indicator from the recording data to be identified.
[0079] In a preferred embodiment of the present invention, after obtaining the user's voice or the device's broadcast voice, feature extraction is performed by speech segmentation, which can reduce the complexity of the model; a common lightweight convolutional model, such as the PP-ASR model, is used to decode the speech features to obtain the final string information generated from the speech.
[0080] S4. Save all physiological indicators and their corresponding physiological data to the preset health management database.
[0081] Preferably, before saving the various physiological indicators and their corresponding physiological data to the preset health management database, the method further includes:
[0082] S41. Detect whether there are non-numeric symbols in each physiological data, and if it is confirmed that there are non-numeric symbols in any physiological data, the data extraction is repeated;
[0083] S42. Detect whether there is a decimal point in each physiological data, and when it is confirmed that there is a decimal point in any physiological data, process the physiological data into integers;
[0084] S43. Detect whether there are numerical symbols in each physiological data, and delete the numerical symbols when it is confirmed that there are numerical symbols in any physiological data.
[0085] In a preferred embodiment of the present invention, in order to ensure the accuracy of data entry, the identified physiological data is subjected to digit symbol detection and non-digit symbol detection.
[0086] Preferably, the step of saving various physiological indicators and their corresponding physiological data into a preset health management database includes:
[0087] S44. Match each physiological indicator with the physiological indicator fields in the health management data table, fill the successfully matched physiological data into the corresponding physiological indicator fields, generate a complete health management data table, and save it to the health management database.
[0088] In a preferred embodiment of the present invention, based on the various physiological indicator fields in the health management data table to be filled established in step S0, the physiological data of the identified physiological indicators are filled into the corresponding physiological indicator fields by a cyclic matching method to generate a complete health management data table and save it to the health management database.
[0089] This embodiment provides a multimodal data input method. It acquires a data input mode; when the data input mode is determined to be an image input mode, it acquires physiological indicator data images and extracts data from these images using a preset image detection model, then outputs various physiological indicators and their corresponding physiological data, and saves them to a preset health management database; when the data input mode is determined to be a voice input mode, it acquires audio recording data and extracts data from these recordings using a preset voice detection model, then outputs various physiological indicators and their corresponding physiological data, and saves them to a preset health management database. Therefore, this invention can quickly acquire various physiological indicator data and save them to a health management database by recognizing images or audio recordings related to physiological indicator data, avoiding the drawbacks of manual input and improving the ease of use of smart health and elderly care products.
[0090] Referring to Figure 2, it is a schematic diagram of the structure of a multimodal data input device provided in an embodiment of the present invention, including:
[0091] The mode selection module is used to acquire the data input mode;
[0092] The image detection module is used to acquire physiological indicator data images when the data input mode is determined to be image input mode, and to extract data from the physiological indicator data images through a preset image detection model, and then output various physiological indicators and the physiological data corresponding to each physiological indicator.
[0093] The voice detection module is used to acquire recording data when the data input mode is determined to be voice input mode, and to extract data from the recording data through a preset voice detection model, and then output various physiological indicators and the physiological data corresponding to each physiological indicator.
[0094] The data storage module is used to save various physiological indicators and their corresponding physiological data to a preset health management database.
[0095] This embodiment provides a multimodal data input device. By acquiring a data input mode, when the data input mode is determined to be an image input mode, physiological indicator data images are acquired, and data is extracted from the physiological indicator data images using a preset image detection model. Then, various physiological indicators and their corresponding physiological data are output and saved to a preset health management database. When the data input mode is determined to be a voice input mode, audio recording data is acquired, and data is extracted from the audio recording data using a preset voice detection model. Then, various physiological indicators and their corresponding physiological data are output and saved to a preset health management database. Therefore, this invention can quickly acquire various physiological indicator data and save them to a health management database by recognizing images or audio recordings related to physiological indicator data, avoiding the drawbacks of manual input and improving the ease of use of smart health and elderly care products.
[0096] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0097] Those skilled in the art will clearly understand that, for convenience and brevity, the specific working process of the device described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0098] Another preferred embodiment of the present invention provides a terminal device including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement a multimodal data input method as described in any of the foregoing embodiments.
[0099] The terminal device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.
[0100] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting all parts of the terminal device via various interfaces and lines.
[0101] The memory can be used to store the computer program. The processor implements various functions of the terminal device by running or executing the computer program stored in the memory and calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc.; the data storage area may store data created based on the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart memory card (SMC), secure digital card (SD), flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0102] Another preferred embodiment of the present invention provides a storage medium, which is a computer-readable storage medium. A computer program is stored in the computer-readable storage medium, and when executed by a processor, the computer program can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0103] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A multimodal data input method, characterized in that, include: Get the data input pattern; When the data input mode is determined to be image input mode, physiological indicator data images are acquired, and data is extracted from the physiological indicator data images through a preset image detection model, and then various physiological indicators and the physiological data corresponding to each physiological indicator are output. When the data input mode is determined to be voice input mode, the recording data is acquired, and the recording data is extracted through a preset voice detection model, and then various physiological indicators and the physiological data corresponding to each physiological indicator are output. All physiological indicators and their corresponding physiological data are saved to a pre-set health management database.
2. The multimodal data input method as described in claim 1, characterized in that, Before obtaining the data input pattern, the following is also included: Obtain several physiological indicator fields and construct a health management data table to be populated.
3. The multimodal data input method as described in claim 2, characterized in that, The step of extracting data from the physiological indicator data image using a preset image detection model, and then outputting various physiological indicators and their corresponding physiological data, includes: Right-angle edge detection is performed on the physiological indicator data image to obtain the corner points of the physiological indicator data image; Based on the corner points, the physiological indicator data image is segmented to generate several segmented images; The segmented image is subjected to perspective transformation to generate several images to be detected; The image to be detected is input into a preset image detection model so that the image detection model can identify and output various physiological indicators and the physiological data corresponding to each physiological indicator.
4. The multimodal data input method as described in claim 3, characterized in that, The step of performing right-angle edge detection on the physiological indicator data image to obtain the corner points of the physiological indicator data image includes: Edge detection is performed on the physiological indicator data image to obtain several edge pixels of the physiological indicator data image; Line fitting is performed on a number of the aforementioned edge pixels to generate a number of line clusters in the physiological index data image; The coordinates of the aforementioned line clusters are transformed to generate the polar coordinates of each line cluster; Based on the polar coordinates of each line cluster, the edge lines of the physiological indicator data image are determined, and based on the edge lines, the corner points of the physiological indicator data image are determined.
5. The multimodal data input method as described in claim 1, characterized in that, The process of extracting data from the recording data using a preset speech detection model, and then outputting various physiological indicators and their corresponding physiological data, includes: The audio data is denoised to generate audio data to be identified; The recording data to be identified is input into the speech detection model so that the speech detection model can extract and output various physiological indicators and the corresponding physiological data from the recording data to be identified.
6. The multimodal data input method as described in claim 1, characterized in that, Before saving the various physiological indicators and their corresponding physiological data to the preset health management database, the following steps are also included: The system detects whether there are non-numeric symbols in each physiological data point, and if any non-numeric symbol is confirmed to exist in any physiological data point, the data is extracted again. The system detects whether there is a decimal point in each physiological data point, and if a decimal point is confirmed in any physiological data point, the physiological data is processed into an integer. The system detects whether there are numerical symbols in each physiological data point, and deletes the numerical symbols when they are confirmed to exist in any physiological data point.
7. The multimodal data input method as described in claim 2, characterized in that, The step of saving various physiological indicators and their corresponding physiological data into a preset health management database includes: The various physiological indicators are matched with the physiological indicator fields in the health management data table, and the successfully matched physiological data is filled into the corresponding physiological indicator fields to generate a complete health management data table, which is then saved to the health management database.
8. A multimodal data input device, characterized in that, include: The mode selection module is used to acquire the data input mode; The image detection module is used to acquire physiological indicator data images when the data input mode is determined to be an image input mode, and to perform data extraction on the physiological indicator data images using a preset image detection model. It retrieves and then outputs various physiological indicators and the corresponding physiological data for each physiological indicator; The voice detection module is used to acquire recording data when the data input mode is determined to be voice input mode, and to extract data from the recording data through a preset voice detection model, and then output various physiological indicators and the physiological data corresponding to each physiological indicator. The data storage module is used to save various physiological indicators and their corresponding physiological data to a preset health management database.
9. A terminal device, characterized in that, The system includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements a multimodal data input method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device where the storage medium is located to perform a multimodal data input method as described in any one of claims 1 to 7.