A multi-modal data input method, device, terminal equipment and storage medium

By identifying and extracting images or recordings related to physiological index data, the problem of manual input in smart health care products is solved, and the rapid and accurate input of data is achieved, and the convenience of using the product is improved.

CN118333563BActive Publication Date: 2025-05-16SHENZHEN TOPTECH MANUFACTORING CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410549556.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-06
Publication Date
2025-05-16
Estimated Expiration
2044-05-06

AI Technical Summary

Technical Problem

Existing smart health care products require manual input of physiological indicator data, which is cumbersome and prone to numerical input errors, resulting in poor usage and difficulty in promotion.

Method used

By identifying images or recordings related to physiological index data, the physiological index data can be quickly extracted using a preset image detection model or speech detection model and saved to the health management database.

Benefits of technology

It realizes the rapid input and accurate storage of physiological index data, avoids the disadvantages of manual input, and improves the convenience of using smart healthy elderly care products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118333563B_ABST
    Figure CN118333563B_ABST
Patent Text Reader

Abstract

The present invention discloses a multimodal data input method, device, terminal equipment and storage medium, which obtains a data input mode; when the data input mode is determined to be an image input mode, the physiological indicator data image is extracted by a preset image detection model, and then various physiological indicators and physiological data corresponding to each physiological indicator are output, and saved in a preset health management database; when the data input mode is determined to be a voice input mode, the recording data is extracted by a preset voice detection model, and then various physiological indicators and physiological data corresponding to each physiological indicator are output, and saved in a preset health management database. Therefore, the present invention can quickly obtain various physiological indicator data and save them in a health management database by identifying images or recordings related to physiological indicator data, avoiding various disadvantages of manual input, and improving the convenience of use of smart health and elderly care products.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing and speech processing, and in particular to a multimodal data input method, device, terminal equipment and storage medium. Background Art

[0002] With the current trend of population structure change in my country, more and more smart products will enter the home-based elderly care scene. How to provide home companionship for the elderly and provide corresponding smart services is a topic of common concern for the current society and enterprises. Smart health care products are generally smart detection equipment, ECG monitoring equipment and blood pressure monitoring equipment, etc. These products can meet the urgent needs and health detection of the elderly in a short period of time. With the development of the new generation of artificial intelligence, new smart health care products can rely on the new generation of artificial intelligence models and other technologies to conduct multi-dimensional analysis such as data mining on various physiological indicators of the elderly, and realize potential disease risk assessment and early intervention.

[0003] However, the methods used by new smart health care products to obtain physiological indicators of the elderly mostly require manual input. The data of physiological indicators is very large, the operation is cumbersome, and manual input often results in numerical input errors. As a result, the user experience of smart health care products is very poor, which ultimately leads to the inability to promote smart health care products.

[0004] Therefore, how to realize the quick input of data for smart health care and elderly care products has become an urgent problem that needs to be solved. Summary of the invention

[0005] The embodiments of the present invention provide a multimodal data input method, apparatus, terminal device and storage medium, which can quickly obtain various physiological indicator data and save them to a health management database by identifying images or recordings related to physiological indicator data, thereby avoiding various disadvantages of requiring manual input and improving the ease of use of smart health care products.

[0006] An embodiment of the present invention provides a multimodal data input method, comprising:

[0007] Get data input mode;

[0008] When it is determined that the data input mode is the image input mode, a physiological indicator data image is obtained, and data is extracted from the physiological indicator data image through a preset image detection model, and then various physiological indicators and physiological data corresponding to the various physiological indicators are output;

[0009] When it is determined that the data input mode is the voice input mode, recording data is obtained, and data extraction is performed on the recording data through a preset voice detection model, and then various physiological indicators and physiological data corresponding to the various physiological indicators are output;

[0010] The various physiological indicators and the physiological data corresponding to the various physiological indicators are saved in a preset health management database.

[0011] Furthermore, before obtaining the data input mode, it also includes:

[0012] Obtain several physiological indicator fields and construct a health management data table to be filled.

[0013] Furthermore, the method of extracting data from the physiological indicator data image by using a preset image detection model, and then outputting various physiological indicators and physiological data corresponding to each physiological indicator, includes:

[0014] Performing right-angle edge detection on the physiological indicator data image to obtain corner points of the physiological indicator data image;

[0015] According to the corner points, the physiological index data image is segmented to generate a plurality of segmented images;

[0016] Performing perspective transformation on the segmented image to generate a plurality of images to be detected;

[0017] The image to be detected is input into a preset image detection model so that the image detection model recognizes and outputs various physiological indicators and physiological data corresponding to the various physiological indicators.

[0018] Furthermore, the performing right-angle edge detection on the physiological indicator data image to obtain corner points of the physiological indicator data image includes:

[0019] Performing edge detection on the physiological indicator data image to obtain a number of edge pixel points of the physiological indicator data image;

[0020] Performing straight line fitting on a plurality of edge pixel points to generate a plurality of straight line clusters of the physiological index data image;

[0021] Performing coordinate transformation on the plurality of straight line clusters to generate polar coordinates of each straight line cluster;

[0022] According to the polar coordinates of each straight line cluster, the edge straight line of the physiological indicator data image is determined, and according to the edge straight line, the corner point of the physiological indicator data image is determined.

[0023] Furthermore, the recording data is extracted by using a preset voice detection model, and then various physiological indicators and physiological data corresponding to the various physiological indicators are output, including:

[0024] De-noising the recorded data to generate recorded data to be identified;

[0025] The recorded data to be recognized is input into the speech detection model, so that the speech detection model extracts and outputs various physiological indicators and physiological data corresponding to the various physiological indicators from the recorded data to be recognized.

[0026] Furthermore, before saving each physiological indicator and the physiological data corresponding to each physiological indicator into a preset health management database, it also includes:

[0027] Detect whether there are non-numeric symbols in each physiological data, and re-extract the data when it is confirmed that there are non-numeric symbols in any physiological data;

[0028] Detecting whether there is a decimal point in each physiological data, and performing integer processing on the physiological data when it is confirmed that there is a decimal point in any physiological data;

[0029] Detect whether there is a digital symbol in each physiological data, and delete the digital symbol when it is confirmed that there is a digital symbol in any physiological data.

[0030] Furthermore, the step of storing each physiological indicator and the physiological data corresponding to each physiological indicator in a preset health management database includes:

[0031] Each physiological indicator is matched with each physiological indicator field in the health management data table, and the successfully matched physiological data is filled into the corresponding physiological indicator field to generate a complete health management data table, and saved in the health management database.

[0032] Another embodiment of the present invention provides a multimodal data input device, comprising:

[0033] Mode selection module, used to obtain data input mode;

[0034] An image detection module, for obtaining a physiological indicator data image when determining that the data input mode is an image entry mode, and performing data extraction on the physiological indicator data image through a preset image detection model, and then outputting various physiological indicators and physiological data corresponding to various physiological indicators;

[0035] A voice detection module, for obtaining recording data when determining that the data input mode is a voice input mode, and extracting data from the recording data through a preset voice detection model, and then outputting various physiological indicators and physiological data corresponding to the various physiological indicators;

[0036] The data saving module is used to save various physiological indicators and the physiological data corresponding to each physiological indicator into a preset health management database.

[0037] Another embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, a multimodal data input method as described in any one of the embodiments is implemented.

[0038] Another embodiment of the present invention provides a storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute a multimodal data input method as described in any one of the above embodiments.

[0039] The following beneficial effects are achieved by implementing the present invention:

[0040] The present invention discloses a multimodal data input method, device, terminal equipment and storage medium, wherein the method acquires a data input mode; when the data input mode is determined to be an image input mode, acquires a physiological indicator data image, and extracts data from the physiological indicator data image through a preset image detection model, and then outputs various physiological indicators and physiological data corresponding to each physiological indicator, and saves them in a preset health management database; when the data input mode is determined to be a voice input mode, acquires recording data, and extracts data from the recording data through a preset voice detection model, and then outputs various physiological indicators and physiological data corresponding to each physiological indicator, and saves them in a preset health management database. Therefore, the present invention can quickly acquire various physiological indicator data and save them in a health management database by identifying images or recordings related to physiological indicator data, thereby avoiding various drawbacks of requiring manual input and improving the ease of use of smart health and elderly care products. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a flowchart of a multimodal data input method provided by an embodiment of the present invention.

[0042] Figure 2 It is a structural schematic diagram of a multimodal data input device provided by one embodiment of the present invention.

[0043] Figure 3It is a flow chart of a method using image recognition input provided by an embodiment of the present invention.

[0044] Figure 4 It is a flow chart of a method using voice recognition input provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by technicians in the technical field to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" in the specification and claims of this application and the above-mentioned figure descriptions and any variations thereof are intended to cover non-exclusive inclusions.

[0047] In the description of the embodiments of the present application, the technical terms "first", "second", etc. are only used to distinguish different objects, and cannot be understood as indicating or implying relative importance or implicitly indicating the number, specific order or primary and secondary relationship of the indicated technical features. In the description of the embodiments of the present application, the meaning of "multiple" is more than two, unless otherwise clearly and specifically defined.

[0048] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0049] In the description of the embodiments of the present application, the term "and / or" is only a description of the association relationship of the associated objects, indicating that there may be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0050] In the description of the embodiments of the present application, the term "multiple" refers to more than two (including two). Similarly, "multiple groups" refers to more than two groups (including two groups), and "multiple pieces" refers to more than two pieces (including two pieces).

[0051] In the description of the embodiments of the present application, unless otherwise clearly specified and limited, technical terms such as "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the internal connection of two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present application can be understood according to the specific circumstances.

[0052] See also Figure 1 , is a flowchart of a multimodal data input method provided by an embodiment of the present invention, comprising:

[0053] S1, obtain data input mode;

[0054] In a preferred embodiment of the present invention, the data input mode is obtained according to the user's operation on the interactive interface of the system. It should be noted that in this embodiment, the data input mode includes: image entry mode, voice entry mode, and manual entry mode.

[0055] Preferably, before obtaining the data input mode, the method further includes:

[0056] S0. Obtain several physiological indicator fields and construct a health management data table to be filled.

[0057] In a preferred embodiment of the present invention, according to the user's operation on the interactive interface of the system, several physiological indicator fields are obtained and a health management data table to be filled is constructed. It is understandable that the user can select several physiological indicator fields that need to be entered this time on the interactive interface to construct the health management data table to be filled.

[0058] S2. When it is determined that the data input mode is the image input mode, a physiological indicator data image is obtained, and data is extracted from the physiological indicator data image through a preset image detection model, and then various physiological indicators and physiological data corresponding to the various physiological indicators are output;

[0059] In a preferred embodiment of the present invention, the camera device carried by the system is turned on to shoot the physiological index data and obtain the physiological index data image. Specifically, the user needs to lay the physical examination report or the physiological index print paper flat in the field of view of the robot camera, keep it still and not blocked, and the ambient light can meet the robot camera to clearly shoot the target area font. In addition, the user can also directly upload the physiological index data image that has been shot.

[0060] Preferably, the method of extracting data from the physiological indicator data image by using a preset image detection model, and then outputting various physiological indicators and physiological data corresponding to each physiological indicator, includes:

[0061] S21, performing right-angle edge detection on the physiological indicator data image to obtain corner points of the physiological indicator data image;

[0062] Preferably, the performing right-angle edge detection on the physiological indicator data image to obtain corner points of the physiological indicator data image includes:

[0063] S211, performing edge detection on the physiological indicator data image to obtain a number of edge pixel points of the physiological indicator data image;

[0064] S212, performing straight line fitting on a plurality of edge pixel points to generate a plurality of straight line clusters of the physiological index data image;

[0065] S213, performing coordinate transformation on the plurality of straight line clusters to generate polar coordinates of each straight line cluster;

[0066] S214 . Determine edge straight lines of the physiological indicator data image according to the polar coordinates of each straight line cluster, and determine corner points of the physiological indicator data image according to the edge straight lines.

[0067] S22, performing image segmentation on the physiological indicator data image according to the corner points to generate a plurality of segmented images;

[0068] S23, performing perspective transformation on the segmented image to generate a plurality of images to be detected;

[0069] S24, inputting the image to be detected into a preset image detection model, so that the image detection model recognizes and outputs various physiological indicators and physiological data corresponding to the various physiological indicators.

[0070] In a preferred embodiment of the present invention, edge detection is first performed on the physiological index data image to obtain all edge points of the image, and the straight line cluster Y=kX+b of the edge points of the image is converted into the polar coordinate space ρ=xCosθ+ySi nθ. It can be understood that for all points on any straight line in the image, there is a relatively strong signal in the polar coordinate space (ρ, θ), so the pixel coordinates of each point on the straight line are obtained by reverse calculation, and the four edge lines of the target area are obtained. By combining the four binary linear straight line equations of the image, the four intersecting points C1, C2, C3, and C4 are solved.

[0071] Furthermore, the quadrilateral area surrounded by the four point straight lines is segmented and then perspective transformed into a rectangle, where the perspective transformation formula is:

[0072]

[0073] Furthermore, the image to be detected is subjected to OCR detection and recognition. Specifically, in this embodiment, a lightweight deep network model PP-OCRv2 model is used to detect and recognize text in the image, and various physiological indicators and physiological data corresponding to each physiological indicator are obtained.

[0074] S3. When it is determined that the data input mode is the voice input mode, obtaining recording data, and performing data extraction on the recording data through a preset voice detection model, and then outputting various physiological indicators and physiological data corresponding to the various physiological indicators;

[0075] In a preferred embodiment of the present invention, the recording data is obtained by turning on the microphone array to automatically monitor the voice of the user or the detection device in the environment. Specifically, the user selects a relatively quiet environment and turns on the system to start the microphone array to monitor the voice in the environment, that is, the sound of the user reading the physical examination report or the physiological index detection result, or the sound of the human body index detection device broadcasting the measurement result. In addition, the user can also directly upload the recorded data in advance.

[0076] Preferably, the data extraction of the recording data is performed by using a preset voice detection model, and then various physiological indicators and physiological data corresponding to the various physiological indicators are output, including:

[0077] S31, performing noise reduction on the recorded data to generate recorded data to be identified;

[0078] S32, inputting the to-be-recognized recorded data into the speech detection model, so that the speech detection model extracts and outputs various physiological indicators and physiological data corresponding to the various physiological indicators from the to-be-recognized recorded data.

[0079] In a preferred embodiment of the present invention, after obtaining the user's voice or the device broadcast sound, feature extraction is performed by voice segmentation to reduce the complexity of the model; a public lightweight convolutional model, such as the PP-ASR model, is used to decode the voice features to obtain the string information ultimately generated by the voice.

[0080] S4. Save various physiological indicators and physiological data corresponding to various physiological indicators into a preset health management database.

[0081] Preferably, before saving each physiological indicator and the physiological data corresponding to each physiological indicator into a preset health management database, the method further includes:

[0082] S41, detecting whether there are non-numeric symbols in each physiological data, and re-extracting the data when it is confirmed that there are non-numeric symbols in any physiological data;

[0083] S42, detecting whether there is a decimal point in each physiological data, and performing integer processing on the physiological data when it is confirmed that there is a decimal point in any physiological data;

[0084] S43, detecting whether there are digital symbols in each physiological data, and deleting the digital symbols when it is confirmed that there are digital symbols in any physiological data.

[0085] In a preferred embodiment of the present invention, in order to ensure the accuracy of data entry, digital symbol detection and non-digital symbol detection are performed on the recognized physiological data.

[0086] Preferably, the step of saving each physiological indicator and the physiological data corresponding to each physiological indicator into a preset health management database includes:

[0087] S44, matching each physiological indicator with each physiological indicator field in the health management data table, and filling the successfully matched physiological data into the corresponding physiological indicator field, generating a complete health management data table, and saving it to the health management database.

[0088] In a preferred embodiment of the present invention, according to the various physiological indicator fields to be filled in the health management data table to be filled in advance in step S0, a circular matching method is adopted to fill the physiological data of each identified physiological indicator into the corresponding physiological indicator field to generate a complete health management data table and save it to the health management database.

[0089] The present embodiment provides a multimodal data input method, by acquiring a data input mode; when determining that the data input mode is an image input mode, acquiring a physiological indicator data image, and extracting data from the physiological indicator data image through a preset image detection model, and then outputting various physiological indicators and the physiological data corresponding to each physiological indicator, and saving them to a preset health management database; when determining that the data input mode is a voice input mode, acquiring recording data, and extracting data from the recording data through a preset voice detection model, and then outputting various physiological indicators and the physiological data corresponding to each physiological indicator, and saving them to a preset health management database. Therefore, the present invention can quickly acquire various physiological indicator data and save them to a health management database by identifying images or recordings related to physiological indicator data, avoiding various disadvantages of requiring manual input, and improving the ease of use of smart health and elderly care products.

[0090] See also Figure 2, is a schematic diagram of the structure of a multimodal data input device provided by an embodiment of the present invention, comprising:

[0091] Mode selection module, used to obtain data input mode;

[0092] An image detection module, for obtaining a physiological indicator data image when determining that the data input mode is an image entry mode, and performing data extraction on the physiological indicator data image through a preset image detection model, and then outputting various physiological indicators and physiological data corresponding to various physiological indicators;

[0093] A voice detection module, for obtaining recording data when determining that the data input mode is a voice input mode, and extracting data from the recording data through a preset voice detection model, and then outputting various physiological indicators and physiological data corresponding to the various physiological indicators;

[0094] The data saving module is used to save various physiological indicators and the physiological data corresponding to each physiological indicator into a preset health management database.

[0095] The present embodiment provides a multimodal data input device, which obtains a data input mode; when it is determined that the data input mode is an image input mode, obtains a physiological indicator data image, and extracts data from the physiological indicator data image through a preset image detection model, and then outputs various physiological indicators and the physiological data corresponding to each physiological indicator, and saves them in a preset health management database; when it is determined that the data input mode is a voice input mode, obtains recording data, and extracts data from the recording data through a preset voice detection model, and then outputs various physiological indicators and the physiological data corresponding to each physiological indicator, and saves them in a preset health management database. Therefore, the present invention can quickly obtain various physiological indicator data and save them in a health management database by identifying images or recordings related to physiological indicator data, avoiding various disadvantages of manual input and improving the ease of use of smart health and elderly care products.

[0096] It should be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the accompanying drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art may understand and implement it without paying any creative effort.

[0097] Those skilled in the art can clearly understand that for the sake of convenience and brevity, the specific working process of the device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0098] Another preferred embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, a multimodal data input method as described in any one of the above embodiments is implemented.

[0099] The terminal device may be a computing device such as a desktop computer, a notebook, a PDA, a cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.

[0100] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the terminal device, and uses various interfaces and lines to connect various parts of the entire terminal device.

[0101] The memory can be used to store the computer program, and the processor realizes various functions of the terminal device by running or executing the computer program stored in the memory and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Med i aCard, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0102] Another preferred embodiment of the present invention provides a storage medium, the storage medium is a computer-readable storage medium, the computer program is stored in the computer-readable storage medium, and the computer program, when executed by the processor, can implement the steps of the above-mentioned various method embodiments. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal and software distribution medium, etc.

[0103] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A multimodal data input method, characterized in that: include: Get data input mode; When it is determined that the data input mode is the image input mode, acquiring a physiological indicator data image; Performing edge detection on the physiological indicator data image to obtain a number of edge pixel points of the physiological indicator data image; Performing straight line fitting on a plurality of edge pixel points to generate a plurality of straight line clusters of the physiological index data image; Performing coordinate transformation on the plurality of straight line clusters to generate polar coordinates of each straight line cluster; Determine edge straight lines of the physiological indicator data image according to the polar coordinates of each straight line cluster, and determine corner points of the physiological indicator data image according to the edge straight lines; According to the corner points, the physiological index data image is segmented to generate a plurality of segmented images; Performing perspective transformation on the segmented image to generate a plurality of images to be detected; Inputting the image to be detected into a preset image detection model so that the image detection model recognizes and outputs various physiological indicators and physiological data corresponding to the various physiological indicators; When it is determined that the data input mode is the voice input mode, recording data is obtained, and data extraction is performed on the recording data through a preset voice detection model, and then various physiological indicators and physiological data corresponding to the various physiological indicators are output; The various physiological indicators and the physiological data corresponding to the various physiological indicators are saved in a preset health management database.

2. A multimodal data input method as claimed in claim 1, characterized in that: Before getting the data input mode, also include: Obtain several physiological indicator fields and construct a health management data table to be filled.

3. A multimodal data input method as claimed in claim 1, characterized in that: The recording data is extracted by using a preset voice detection model, and then various physiological indicators and physiological data corresponding to the various physiological indicators are output, including: De-noising the recorded data to generate recorded data to be identified; The recorded data to be recognized is input into the speech detection model, so that the speech detection model extracts and outputs various physiological indicators and physiological data corresponding to the various physiological indicators from the recorded data to be recognized.

4. A multimodal data input method as claimed in claim 1, characterized in that: Before saving each physiological indicator and the physiological data corresponding to each physiological indicator into a preset health management database, it also includes: Detect whether there are non-numeric symbols in each physiological data, and re-extract the data when it is confirmed that there are non-numeric symbols in any physiological data; Detecting whether there is a decimal point in each physiological data, and performing integer processing on the physiological data when it is confirmed that there is a decimal point in any physiological data; Detect whether there is a digital symbol in each physiological data, and delete the digital symbol when it is confirmed that there is a digital symbol in any physiological data.

5. A multimodal data input method as claimed in claim 2, characterized in that: The step of storing various physiological indicators and physiological data corresponding to the various physiological indicators in a preset health management database includes: Each physiological indicator is matched with each physiological indicator field in the health management data table, and the successfully matched physiological data is filled into the corresponding physiological indicator field to generate a complete health management data table, and saved in the health management database.

6. A multimodal data input device, characterized in that: include: Mode selection module, used to obtain data input mode; An image detection module is used to obtain a physiological indicator data image when determining that the data input mode is an image entry mode, perform edge detection on the physiological indicator data image, and obtain a number of edge pixel points of the physiological indicator data image; perform straight line fitting on a number of the edge pixel points to generate a number of straight line clusters of the physiological indicator data image; perform coordinate conversion on a number of the straight line clusters to generate polar coordinates of each straight line cluster; determine the edge straight lines of the physiological indicator data image according to the polar coordinates of each straight line cluster, and determine the corner points of the physiological indicator data image according to the edge straight lines; perform image segmentation on the physiological indicator data image according to the corner points to generate a number of segmented images; perform perspective transformation on the segmented images to generate a number of images to be detected; input the image to be detected into a preset image detection model so that the image detection model recognizes and outputs various physiological indicators and physiological data corresponding to each physiological indicator; A voice detection module, for obtaining recording data when determining that the data input mode is a voice input mode, and extracting data from the recording data through a preset voice detection model, and then outputting various physiological indicators and physiological data corresponding to the various physiological indicators; The data saving module is used to save various physiological indicators and the physiological data corresponding to each physiological indicator into a preset health management database.

7. A terminal device, characterized in that: It comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, a multimodal data input method as described in any one of claims 1 to 5 is implemented.

8. A storage medium, characterized in that: The storage medium includes a stored computer program, wherein when the computer program is executed, the device where the storage medium is located is controlled to execute a multimodal data input method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Multi-link data collecting method for document shorthand book

    CN108008824A

  • Image structured data extraction method, electronic device and storage medium

    CN111695439A

  • Anti-noise mobile inspection voice semantic recognition system

    CN114171017A

  • Visual measurement method and system for size of steel plate

    CN114494030A