Multimodal data integration and augmentation system, and data processing method therefor
The system addresses the challenges of integrating and analyzing multimodal healthcare data by using a comprehensive data processing framework that improves data quality and integrates it into a time series, reducing computational load and enhancing analysis accuracy.
Patent Information
- Application Number
- PCT/KR2024/017794
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-24
- Filing Date
- 2024-11-12
- Publication Date
- 2025-05-30
AI Technical Summary
Existing data integration systems in healthcare face challenges in efficiently processing and analyzing multimodal data from various devices, particularly due to issues with missing or inaccurate data, and the inability to integrate data into a time series effectively, leading to a high computational load on artificial intelligence systems.
A system that integrates multimodal data by using a data acquisition unit with various devices, a data processing unit that classifies and improves data quality, and a data study unit that analyzes the integrated data using artificial intelligence. The system includes units for managing, evaluating, and augmenting data to ensure high-quality, time-series integrated data.
The system enables efficient analysis of multimodal data by improving data quality, integrating data into a time series, and reducing the computational load on AI systems, thereby enhancing the accuracy and reliability of healthcare services.
Smart Images

Figure KR2024017794_30052025_PF_FP_ABST
Abstract
Description
A system for integrating and augmenting multimodal data and a data processing method thereof
[0001] The present invention relates to a system that integrates and enhances multimodal data input through multiple input devices, and more specifically, to a system that integrates video, images, voice, surrounding environment, and biometric data for an individual profile in a time series manner, and improves data with insufficient quality and creates and integrates new data.
[0002]
[0003] In general, technologies are being developed to measure users' health status based on various types of information obtained through various devices in various fields, including the healthcare field, to provide information to users.
[0004] This technology integrates various sensors and data to assess a user's health from various perspectives. For example, video data can assess exercise posture, physical condition, and the surrounding environment, while biometric sensors can provide physiological data such as heart rate and blood pressure.
[0005] Additionally, it is possible to conduct a comprehensive health assessment, such as analyzing the characteristics of voice and ambient noise to understand emotional aspects.
[0006] Through various data sources, personalized health management information can be provided to individual users. For example, a user's exercise patterns can be analyzed using video data to suggest areas for improvement, or biometric sensor data can be utilized to recommend diet and exercise programs tailored to the individual's health status.
[0007] Furthermore, analyzing aggregated data can detect early signs of disease and provide information that aids in prevention. For example, biometric sensor data can detect changes in heart rate or blood pressure, allowing for early detection of abnormalities.
[0008] Image data can be used to visually assess a patient's condition, which can aid in diagnosis and treatment. Furthermore, data on voice and ambient sounds can be collected to improve voice diagnosis and communication with patients.
[0009] By integrating diverse data in this way, we can enable comprehensive and customized healthcare services, helping users manage their health more effectively.
[0010] Healthcare-related data integration systems hold significant growth potential, and technological advancements and innovations, coupled with advancements in artificial intelligence, will enable more sophisticated data analysis and interpretation.
[0011] Furthermore, advancements in artificial intelligence can provide additional information that can aid healthcare, using indirect data that is currently unavailable. This enhanced ability to analyze diverse types of data and transform them into useful information can increase the growth potential of related fields.
[0012] Additionally, as interest in health management increases, interest in and demand for health are increasing along with changes in population structure, and more people can use healthcare services to maintain health and prevent disease.
[0013] Personalized healthcare services are expected to have significant growth potential for companies that provide them, as users are increasingly seeking personalized health services and advice tailored to their individual needs.
[0014] Furthermore, convergence with the medical industry can lead to large-scale innovations in diagnosis, treatment, and monitoring. As the medical industry is becoming increasingly digitalized, the importance of related technologies is increasing.
[0015] However, simply inputting each piece of data into artificial intelligence by acquiring various pieces of information from different devices places a significant burden on the artificial intelligence, so a technology that can reasonably integrate each piece of data is needed.
[0016] To solve this problem, U.S. Patent No. 10-2233725 discloses a 5G-based multi-healthcare service smart home platform system and its operating method.
[0017] However, such a technique may incur a large load in computing data because it cannot determine missing parts in the acquired data and cannot identify differences between different data.
[0018] In addition, according to Korean Patent No. 2188766, the user's daily exercise amount
[0019] We are introducing an artificial intelligence-based healthcare service provider that can provide information and services related to improving the user's physical functions and potential diseases based on the information.
[0020] However, this configuration may not address issues that may arise when analyzing data containing missing or inaccurate data. Furthermore, because individual data cannot be integrated into a time series and provided to AI, the AI may incur significant overhead in calculating each piece of data.
[0021]
[0022] The present invention has been devised based on the above technical background, and provides a system that generates multimodal data that is easy to analyze by integrating individual user profile data acquired through a plurality of different devices into multimodal data, integrating each data into a time series, and creating new parts that are lacking in quality or missing for each data, or processing the data to integrate high-quality data.
[0023]
[0024] In order to achieve the above purpose, the present invention includes a data acquisition unit including a photographing device, a recording device, an environmental sensor device, and a biosensor device, a data processing unit that receives each data acquired from the data acquisition unit and generates integrated data, and a data study unit that analyzes the integrated data through artificial intelligence to derive a result value, so that the data processing unit can judge the quality of each data acquired from the data acquisition unit and improve the quality of each data.
[0025] In addition, the data processing unit may include a data management unit that classifies a plurality of different data provided from the data acquisition unit by type, a missing value confirmation unit that determines missing values for each of the classified plurality of data, a missing value processing unit that processes each data according to the missing values and generates processing data, a data evaluation unit that evaluates the quality of the processing data and determines low-quality data that is difficult to analyze through artificial intelligence and analyzable usage data, a data enhancement unit that enhances the low-quality data to improve its quality and generates improved data, and a data integration unit that integrates all of the improved data and the usage data generated according to each of the data classified by type to generate integrated data.
[0026] In addition, the data management unit can classify the plurality of data acquired from the data acquisition unit into image data, audio data, and text data according to their formats.
[0027] In addition, the missing value confirmation unit can classify data with less than 10% of missing values for each classified data as good data, classify data with more than 10% but less than 50% of missing values as supplementary data, and classify data with more than 50% of missing values as discarded data.
[0028] In addition, the missing data processing unit can delete missing portions from the good data, replace the missing portions of the supplementary data with surrounding values through the KNN algorithm in the supplementary data, and delete the entire discarded data.
[0029] In addition, the data evaluation unit can perform grammatical error checks, local consistency checks, global consistency checks, and readability checks on the text data.
[0030] In addition, the above grammatical error check can be confirmed through the grammatical error inclusion rate obtained through the following [Mathematical Formula 1].
[0031] [Mathematical Formula 1]
[0032]
[0033] In addition, the above local consistency test determines the relevance of adjacent different sentences, and the relevance can be confirmed through the following [Mathematical Formula 2].
[0034] [Equation 2]
[0035]
[0036] In addition, the global consistency check can analyze the relationship between each sentence through the following [Mathematical Formula 3] to determine whether all sentences that make up the entire text data are related to one topic.
[0037] [Equation 3]
[0038]
[0039] In addition, the above interrelationship can be determined by determining whether the same word appears in different sentences through [Mathematical Formula 4], and the interrelationship between different sentences can be corrected through distance weighting according to the position of the same word appearing in different sentences.
[0040] [Equation 4]
[0041]
[0042] In addition, the readability test is measured through the average length of sentences included in the text data, and the length of sentences and the length of paragraphs included in the text data can be evaluated through the following [Mathematical Formula 5].
[0043] [Equation 5]
[0044]
[0045] In addition, the audio data can be evaluated by the data evaluation unit by converting the signal-to-noise ratio included in the audio data into a spectrogram.
[0046] In addition, the image data can be used to determine the quality of the image through the degree of sharpness and blurring of the image, and the sharpness can be detected by detecting the degree of hue and saturation using the HSV color model.
[0047] In addition, the data augmentation unit can generate and improve data through a deep neural network if the low-quality data is in the form of text or image, and can generate the improved data by improving and generating the low-quality data in the form of audio by converting the low-quality data in the form of audio into an image and then generating the same data as the low-quality data in the form of image.
[0048] In addition, the data integration unit further includes time series information in the usage data and the improvement data formed by augmenting each data classified by type, and each of the data can generate integrated data by matching the cycle and period in which each data is created based on the time series information included in the usage data and the improvement data.
[0049] In addition, a server is formed that transmits external data to the data management unit to supplement each data measured in the data acquisition unit, and the external data can be classified according to the form in which the external data is composed in the data acquisition unit.
[0050] In addition, each piece of data measured in the data acquisition unit is stored to form a database that creates existing data, and existing data measured in the past is transferred to the data management unit, and can be classified according to the form in which the existing data is composed.
[0051]
[0052] A system for integrating and augmenting multimodal data according to one embodiment of the present invention can integrate data acquired through various devices into a single data and derive a result value analyzed through artificial intelligence.
[0053] A system for integrating and augmenting multimodal data according to one embodiment of the present invention can improve the quality of each data and generate one integrated data.
[0054] A system for integrating and augmenting multimodal data according to one embodiment of the present invention can easily process data by classifying a plurality of different data acquired from various devices according to the form of each data.
[0055] A system for integrating and augmenting multimodal data according to one embodiment of the present invention can select an appropriate processing method depending on the proportion of missing values included in the data.
[0056] A system for integrating and augmenting multimodal data according to one embodiment of the present invention can check the quality of text data by checking for grammatical errors, consistency, readability, etc. of text-type data.
[0057] A system for integrating and augmenting multimodal data according to one embodiment of the present invention can evaluate the quality of audio data using a signal-to-noise ratio.
[0058] A system for integrating and augmenting multimodal data according to one embodiment of the present invention can determine the quality of image data through the degree of clarity and blurring of the image data.
[0059] A system for integrating and augmenting multimodal data according to one embodiment of the present invention can improve and generate each piece of data determined to be of low quality through a deep neural network, thereby changing the data into a level at which artificial intelligence can generate a result value.
[0060] A system for integrating and augmenting multimodal data according to one embodiment of the present invention can integrate each piece of data according to a time series, thereby preventing data from being omitted during the integration process.
[0061] A system for integrating and augmenting multimodal data according to one embodiment of the present invention can generate larger-scale data by additionally receiving and integrating information that can supplement information acquired by a data acquisition unit from an external server and database.
[0062]
[0063] FIG. 1 is a configuration diagram of a system for integrating and increasing multimodal data according to one embodiment of the present invention.
[0064] Figure 2 is a detailed diagram of information provided to a data processing unit according to one embodiment of the present invention.
[0065] Figure 3 is a detailed configuration diagram of a data processing unit and a detailed diagram of a data study according to one embodiment of the present invention.
[0066] Figure 4 is an operation diagram of a data management unit according to one embodiment of the present invention.
[0067] Figure 5 is an operation diagram of a missing value confirmation unit according to one embodiment of the present invention.
[0068] Figure 6 is an operation diagram of a missing value processing unit according to one embodiment of the present invention.
[0069] Figure 7 is a diagram illustrating an evaluation process for text data of a data evaluation unit according to one embodiment of the present invention.
[0070] Figure 8 is a diagram illustrating an evaluation process for image data of a data evaluation unit according to one embodiment of the present invention.
[0071] FIG. 9 is a diagram illustrating audio data according to one embodiment of the present invention.
[0072] Figure 10 is a result diagram excluding the erection element of image data according to one embodiment of the present invention.
[0073] FIG. 11 is a diagram of color tone and saturation information of image data according to one embodiment of the present invention.
[0074] Figure 12 is a result of extracting blur from image data according to one embodiment of the present invention.
[0075] Figure 13 is a diagram of the configuration of data provided to a data integration unit according to one embodiment of the present invention.
[0076] Figure 14 is a diagram of the integration process of a data integration unit according to one embodiment of the present invention.
[0077] Figure 15 is a process diagram illustrating a time series combination process in a data integration unit according to one embodiment of the present invention.
[0078] Figure 16 is a flowchart of a data processing method of a data processing system according to one embodiment of the present invention.
[0079] Figure 17 is an operational flow diagram of a data processing unit according to one embodiment of the present invention.
[0080]
[0081] Hereinafter, a preferred embodiment of the present invention will be described in detail with reference to the attached drawings.
[0082] The advantages and features of the present invention and the method for achieving them will become clear with reference to the embodiments described in detail below together with the attached drawings.
[0083] However, the present invention is not limited to the embodiments disclosed below, but may be implemented in various different forms, and the present embodiments are provided only to make the disclosure of the present invention complete and to fully inform a person having ordinary skill in the art to which the present invention pertains of the scope of the invention, and the present invention is defined only by the scope of the claims.
[0084] In addition, when describing the present invention, if it is determined that related known technologies or the like may obscure the gist of the present invention, a detailed description thereof will be omitted.
[0085]
[0086] FIG. 1 is a configuration diagram of a system for integrating and increasing multimodal data according to one embodiment of the present invention.
[0087] Figure 2 is a detailed diagram of information provided to a data processing unit according to one embodiment of the present invention.
[0088] Figure 3 is a detailed configuration diagram of a data processing unit and a detailed diagram of a data study according to one embodiment of the present invention.
[0089] Figure 4 is an operation diagram of a data management unit according to one embodiment of the present invention.
[0090] FIG. 5 is an operation diagram of a missing value confirmation unit according to one embodiment of the present invention, and FIG. 6 is an operation diagram of a missing value processing unit according to one embodiment of the present invention.
[0091] FIG. 7 is a diagram showing an evaluation process for text data of a data evaluation unit according to one embodiment of the present invention, and FIG. 8 is a diagram showing an evaluation process for image data of a data evaluation unit according to one embodiment of the present invention.
[0092] FIG. 9 is a diagram illustrating audio data according to one embodiment of the present invention, FIG. 10 is a diagram illustrating the results of excluding the triggering elements of image data according to one embodiment of the present invention, FIG. 11 is a diagram illustrating color tone and saturation information of image data according to one embodiment of the present invention, and FIG. 12 is a diagram illustrating the results of extracting blur from image data according to one embodiment of the present invention.
[0093] FIG. 13 is a diagram of the configuration of data provided to a data integration unit according to one embodiment of the present invention, FIG. 14 is a diagram of an integration process of a data integration unit according to one embodiment of the present invention, and FIG. 15 is a diagram of a process of combining in accordance with a time series in a data integration unit according to one embodiment of the present invention.
[0094] FIG. 16 is a flowchart of a data processing method of a data processing system according to one embodiment of the present invention, and FIG. 17 is an operation flowchart of a data processing unit according to one embodiment of the present invention.
[0095] Referring to FIG. 1, the data processing system (1) can generate photographing data (311), recording data (331), environmental data (351), and biometric data (371) from a data acquisition unit (300) including a photographing device (310), a recording device (330), an environmental sensor device (350), and a biometric sensor device (370), respectively.
[0096] At this time, the above-mentioned shooting device (310) is equipment such as CCTV or a camera, and generates shooting data (311) in the form of video and images, and the recording device (330) is equipment such as a microphone, and generates audio recording data (331) for voice and ambient sounds.
[0097] The above environmental sensor device (350) obtains various environmental data (351) about the surrounding environment, and the biometric sensor device (370) may be biometric data (371) that measures changes in the user's body, including blood pressure, heart rate, body temperature, etc. obtained from a wearable device, etc.
[0098] The environmental data (351) and the biometric data (371) generated from the environmental sensor device (350) and the biometric sensor device (370) can be generated in text form.
[0099] The photographing data (311), the recording data (331), the environmental data (351), and the biometric data (371) generated in this manner may be multimodal data.
[0100] Each piece of data acquired from the data acquisition unit (300) is moved to the data processing unit (100), and each piece of data can be integrated into a time series to generate integrated data (191).
[0101] At this time, the data processing unit (100) can perform preprocessing to classify each piece of information measured by the data acquisition unit (300) according to its form and supplement missing values.
[0102] In addition, the data processing unit (100) can check the quality of data, improve the quality of data, or create new data based on existing data to enhance the data.
[0103] A data study unit (500) that analyzes the integrated data (191) through artificial intelligence to derive a result value may be included, and a data provision unit (700) that provides the result value generated in the data study unit (500) to the user may be formed.
[0104] In addition, the above data study (500) can derive multiple result values from the above integrated data (191).
[0105] Referring to Fig. 2, the data acquisition unit (300) may be formed with multiple devices. The data acquisition unit (300) may include a photographing device (310), a recording device (330), an environmental sensor device (350), and a biosensor device (370).
[0106] The above-mentioned filming device (310) may be a CCTV or a camera that films the user, and the filming device (310) may film the user's actions or provide filming data (311) in the form of video and images of the user's surroundings.
[0107] The above recording device (330) may be a recorder or the like that records the user's voice, and the recording device (330) may provide audio recording data (331) regarding the user's voice and noise around the user.
[0108] The above environmental sensor device (350) is a sensor located in a building or structure and can provide text-type environmental data (351) that can check the surrounding environment of the user's location, such as temperature, humidity, and brightness.
[0109] The above biosensor device (370) includes a wearable device worn by the user, and through this, biometric data (371) that can check the user's oxygen, heart rate, blood pressure, electrocardiogram, and body temperature, etc., can be provided in the form of text.
[0110] Each of the above-described photographing data (311), the above-described recording data (331), the above-described environmental data (351), and the above-described biometric data (371) can be transmitted to the data processing unit (100). The data processing unit (100) can analyze each of the provided information, measure missing values to be described later, improve the quality of each data, and integrate each data according to a time series.
[0111] In addition, the data processing unit (100) can receive external data (910) that can supplement the photographing data (311), the recording data (331), the environmental data (351), and the biometric data (371) through a server (900) located externally.
[0112] External data (910) to supplement each data measured in the above data acquisition unit (300) may be indirect data including direct data.
[0113] For example, the data management unit (110) may provide information on the weather and temperature of the area where the user is located. In addition, to supplement the biometric data (371), information such as blood pressure, heart rate, and body temperature of other people of similar age to the user may be provided.
[0114] In this way, the external data (910) can be provided to supplement the photographing data (311), the recording data (331), the environmental data (351), and the biometric data (371).
[0115] The above external data (910) can be classified according to the form in which each data is composed, together with each data measured by the data acquisition unit (300).
[0116] In addition, each data measured in the data acquisition unit (300) is stored to form a database (800) in which existing data (810) is formed, and existing data (810) measured in the past is transmitted to the data management unit (110) and can be classified according to the form that constitutes the existing data (810).
[0117] At this time, the existing data (810) may be the user's past movements and voice, the environment in which the user was located in the past, and the user's past biometric data (371).
[0118] In this way, a plurality of data are provided to the data processing unit (100), and the data processing unit (100) can integrate each data to generate integrated data (191).
[0119] Referring to FIG. 3, the data processing unit (100) may be formed with a data management unit (110) that classifies a plurality of different data provided from the data acquisition unit (300), server (900), and database (800) by type.
[0120] The above data management unit (110) can classify, in addition to the photographing data (311), recording data (331), environmental data (351) and biometric data (371) acquired from the data acquisition unit (300), external data (910) provided from the server (900) and existing data (810) stored in the database (800), according to the type of each data.
[0121] The above-mentioned shooting data (311), the above-mentioned recording data (331), the above-mentioned environmental data (351), the above-mentioned biometric data (371), the above-mentioned external data (910), and the above-mentioned existing data (810) can be classified into image data (111), audio data (113), and text data (115) according to their form.
[0122] Additionally, the data management unit (110) can store each acquired data in a database (800).
[0123] The above data processing unit (100) may be provided with a missing value confirmation unit (130) that determines missing values for each of the image data (111), the audio data (113), and the text data (115) that are formed by classification.
[0124] In addition, in the above missing value confirmation unit (130), a missing value processing unit (140) may be formed that deletes or processes the missing portion included in each data according to the ratio of missing values included in each data to generate processing data (143).
[0125] A data evaluation unit (160) may be included to evaluate the quality of the above-mentioned processing data (143) and classify the processing data (143) into low-quality data (161) that may be difficult to analyze through artificial intelligence and usable data (163) that can be provided to artificial intelligence for analysis.
[0126] In the case of the above low-quality data (161), since it is unusable data, a separate augmentation process needs to be performed, and therefore, a data augmentation unit (170) that augments the low-quality data (161) to improve quality and generate improved data (171) can be formed.
[0127] It may include a data integration unit (190) that integrates all of the improvement data (171) and usage data (163) generated from each of the above-mentioned image data (111), the above-mentioned audio data (113), and the above-mentioned text data (115) classified by type to generate integrated data (191).
[0128] That is, each of the above image data (111), the above audio data (113) and the above text data (115) can be integrated by using the improvement data (171) and the usage data (163).
[0129] The above integrated data (191) can be provided to data study (500), and the integrated data (191) can be analyzed through artificial intelligence to provide the result value required by the user.
[0130] Referring to FIG. 4, the data management unit (110) can classify the plurality of data acquired from the data acquisition unit (300) into image data (111), audio data (113), and text data (115) according to their formats.
[0131] That is, the data acquisition unit (300) can classify data on the user's life log, such as shooting data (311), recording data (331), environmental data (351), and biometric data (371) collected in daily life, according to the type of each data.
[0132] In this way, by classifying each data generated in various forms such as video, sound, image, and text into the video data (111), the audio data (113), and the text data (115), various data can be appropriately processed according to their forms.
[0133] Additionally, the data management unit (110) can obtain external data (910) through a server (900) located externally. The external data (910) can include various types of data such as video, sound, image, and text.
[0134] Accordingly, by classifying the external data (910) into the image data (111), the audio data (113), and the text data (115), appropriate processing can be performed on the external data (910).
[0135] The existing data (810) stored in the database (800) can be acquired through the above data management unit (110). Since the existing data (810) is past information acquired by the data acquisition unit (300), the existing data (810) can also include various types of data, similar to the photographing data (311), recording data (331), environmental data (351), and biometric data (371) currently acquired by the data acquisition unit (300).
[0136] The above data management unit (110) can classify the existing data (810) into the image data (111), the audio data (113), and the text data (115).
[0137] Referring to FIGS. 5 and 6, the missing value confirmation unit (130) can check the missing values of each image data (111), audio data (113), and text data (115) classified and formed in the data management unit (110).
[0138] It is necessary to check for missing values in video data (111), audio data (113), and text data (115) and process the missing values included in each data. Missing values mean no value or an error value.
[0139] Missing values like these can be difficult to use for general calculations or machine learning tasks.
[0140] In the above missing value confirmation unit (130), the image data (111), audio data (113) and the text data (115) are each confirmed, and if the proportion of missing values is less than 10%, the data can be classified as good data (131).
[0141] If the missing value of the above video data (111), audio data (113) and text data (115) is 10% or more and less than 50%, it can be classified as supplementary data (133), and if the missing value of the above video data (111), audio data (113) and text data (115) is 50% or more, it can be classified as discarded data (135).
[0142] The good data (131), the supplementary data (133), and the discarded data (135) classified in the missing data confirmation unit (130) are moved to the missing data processing unit (140) and can be changed into processed data (143) through different processes.
[0143] The above good data (131) can generate the processing data (143) by deleting the corresponding data record with missing values less than 10% of the total data.
[0144] That is, when generating the processing data (143) from the above good data (131), the processing data (143) can be generated by deleting the missing portion in the above good data.
[0145] The above supplementary data (133) can be used to replace missing values with nearby values through the KNN (K Neighbor Nearest) algorithm in the missing value processing unit (140) to generate processed data (143).
[0146] The above discarded data (135) can generate the processed data (143) by deleting the corresponding column after data review. That is, the above discarded data (135) generates the processed data (143) by deleting all data to which the missing values belong.
[0147] Referring to FIGS. 7 and 8, the quality of the processing data (143) generated by the missing value processing unit (140) can be evaluated. At this time, the quality of the processing data (143) can be evaluated according to the form of the processing data (143).
[0148] That is, the process of evaluating the processed data (143) generated by checking and processing missing values of image data (111) and the processed data (143) generated by checking and processing missing values of text data (115) may be different.
[0149] Figure 7 is a process for evaluating the processed data (143) generated by checking and processing missing values of the above text data (115). The processed data (143) may first undergo a preprocessing process.
[0150] In the above preprocessing step, the processing data (143) may be separated by periods to arrange them into a single sentence. Additionally, if a sentence has more than 50 characters, it may be separated by commas.
[0151] The processed data (143) classified in this manner can be used to create a meaningful morpheme list dictionary and a part-of-speech morpheme frequency dictionary. In addition, duplicate and meaningless sentences can be deleted.
[0152] The above data evaluation unit (160) can perform a grammatical error check (164a), a local consistency check (164b), a global consistency check (164c), and a readability check (164d) on the above text data (115).
[0153] The above grammatical error check (164a) can cause grammatical errors or semantic analysis errors in the text, and thus requires inspection. [Mathematical Formula 1] can be used to calculate the grammatical error inclusion rate for the grammatical error check (164a).
[0154] [Mathematical Formula 1]
[0155]
[0156] Here, is the total number of documents, appeared in the document It can return 1 if the th word is an error, 0 otherwise.
[0157] In addition, the local consistency check (164b) and global consistency check (164c) are for evaluating whether a document in the text data (115) deals with a consistent topic.
[0158] To this end, the degree of association between words appearing in each sentence forming the text data (115) can be calculated. In other words, it is to confirm whether adjacent sentences included in the text data (115) are semantically related to each other.
[0159] The above local consistency check (164b) checks the relevance of adjacent sentences, and the above global consistency check (164c) determines whether all sentences that make up the text data (115) are related to one topic.
[0160] The above local consistency test (164b) determines the relevance of adjacent different sentences, and can be confirmed through [Mathematical Formula 2].
[0161] [Equation 2]
[0162]
[0163] Here, is the number of total sentences, silver The second sentence and It shows the interrelationship between the second sentence.
[0164] In addition, the above global consistency check (164c) can analyze the relationship between each sentence through [Mathematical Formula 3] to determine whether all sentences that make up the entire text data are related to one topic.
[0165] [Equation 3]
[0166]
[0167] Here, is the number of total sentences, silver The second sentence and It shows the interrelationship between the second sentence.
[0168] The interrelationship between the above [Mathematical Formula 2] and the above [Mathematical Formula 3] can be calculated when specific words (nouns, pronouns, noun phrases, etc.) appear simultaneously in different sentences.
[0169] If there is at least one relationship between two different sentences, it indicates the above relationship. The value of can be returned as 1, otherwise it can be calculated as 0.
[0170] The above interrelationship can be determined by determining whether the same word appears in different sentences through [Mathematical Formula 4], and the interrelationship between different sentences can be corrected through distance weighting according to the position of the same word appearing in different sentences.
[0171] [Equation 4]
[0172]
[0173] Here, is a value for the relationship, which returns 1 if at least one identical word appears in another sentence, otherwise it is calculated as 0. Also represents the distance between different sentences in which the above word appears.
[0174] The above readability test (164d) may be a criterion indicating the difficulty of the text data (115). It may be primarily measured by the average length of sentences included in the text data (115). A regression model formula consisting of the total length of the text data (115) and the length of each included sentence may be used.
[0175] The above readability test (164d) can evaluate the length of sentences and paragraphs included in the text data (115) through [Mathematical Formula 5].
[0176] [Equation 5]
[0177]
[0178] Here, is the average paragraph length, is the average sentence length, and the larger the calculated value, the higher the level of the text data can be judged.
[0179] The quality of the processing data (143) generated from the text data (115) can be judged through a text inspection method (164) including a grammatical error check (164a), a local consistency check (164b), a global consistency check (164c), and a readability check (164d), and the data can be classified into usable data (163) that can be analyzed by artificial intelligence and low-quality data (161) that cannot be analyzed by artificial intelligence.
[0180] Figure 8 is a process for evaluating processed data (143) generated by checking and processing missing values of the above image data (111). In order to evaluate the above image data (111), the evaluation can be based on blurring, noise, and color tone changes.
[0181] The above image data (111) can be separated into brightness components through a frequency domain filtering method. The brightness components can be divided into low frequency and high frequency components, and the low frequency components can be lighting components within the image, and the high frequency components can be objects within the image.
[0182] The above image data (111) excluding the brightness element can obtain the purity of hue and saturation through the HSV color model (167). The HSV color model (167) is a color model composed of hue, saturation, and brightness.
[0183] Hue is the degree to which a color appears, and saturation is the degree to which a color is pure. Since brightness factors have been removed, the quality of the image data (111) can be assessed using saturation for hue and clarity.
[0184] Any image that has lost information due to blurring, shaking, ripples, compression, etc., included in the above image data (111) can be recognized as a blurred image. Therefore, the image data (111) can be used to evaluate the image quality through the image clarity and degree of blurring.
[0185] Blur can be measured through the gradient of the gray levels of adjacent pixels, since the more severe the blur, the more likely it is that adjacent pixels will converge to the same gray level.
[0186] In addition, the audio data (113) is converted into a spectrogram (165) and evaluated by using the signal noise ratio (SNR) to image the audio data (113).
[0187] Referring to Figures 9 to 12, Figure 9 is a drawing showing the imaging of the audio data (113). That is, the audio data (113) is converted into a spectrogram (165) and evaluated.
[0188] In addition, in FIG. 10, the image data (111) can exclude brightness elements from the input data through frequency domain filtering.
[0189] The values for hue and saturation can be calculated through the HSV color model (167) in the brightness exclusion drawing (166) generated as a result of excluding the high-frequency part.
[0190] Fig. 11 shows a hue diagram (167a) and a saturation diagram (167b) in which the brightness exclusion diagram (166) is converted to a state in which hue and saturation can be calculated through the HSV color model (167). In addition, Fig. 12 shows a blur diagram (169) in which the brightness exclusion diagram (166) is converted to a state in which blur can be extracted and evaluated through the HSV color model (167).
[0191] By checking the blur, hue, and saturation of the audio data (113) and the image data (111), each processing data (143) generated from the audio data (113) and the image data (111) can be evaluated to distinguish between low-quality data (161) and usage data (163).
[0192] Referring to FIGS. 13 to 15, the data evaluation unit (160) evaluates the quality of each processed data (143) generated by processing missing values of image data (111), audio data (113), and text data (115), thereby classifying the utilized data (163) and low-quality data (161).
[0193] At this time, the usage data (163) may be data that can create a result value that can be provided to the user through artificial intelligence in the data study (500). In addition, the low-quality data (161) may be data that cannot derive a result value through the artificial intelligence due to insufficient quality.
[0194] Since the above low-quality data (161) is difficult to use, a process of enhancing the above low-quality data (161) to generate improved data (171) may be necessary.
[0195] To this end, the data enhancement unit (170) can improve the low-quality data (161) generated for each of the image data (111), the audio data (113) and the text data (115) according to the form of the image data (111), the audio data (113) and the text data (115).
[0196] The process of converting the above low-quality data (161) into the above improved data (171) can be accomplished through quality improvement and creation of new data.
[0197] Quality improvement and generation can be performed on text data (115), audio data (113), and image data (111) according to the form of the above low-quality data (161). Here, audio data (113) is converted into a spectrogram (165) and processed in the same manner as image data (111).
[0198] The quality of the above text data (115) is improved using a transformer encoder. The method applies dropout twice to obtain two different embeddings, which are then improved using positive pairs.
[0199] That is, in unsupervised learning, learning is performed using positive pairs, and in supervised learning, sentences expressing the same meaning are made into positive pairs, and different sentences are made into negative pairs, and the encoder is trained by predicting the positive sentence among the negative sentences.
[0200] The quality of the imaged audio data (113) and video data (111) is improved by using the DnCNN model. In addition, the data can be generated using the generative AI model of DCGAN.
[0201] In this way, when the low-quality data (161) is improved in the data enhancement unit (170) to form improved data, the usage data (163) and the improved data (171) can be moved to the data integration unit (190) to form integrated data (191).
[0202] The data integration unit (190) can receive the usage data (163) and the improvement data (171) included in each of the video data (111), the audio data (113) and the text data (115) and integrate the video data (111), the audio data (113) and the text data (115).
[0203] At this time, the video data (111), the audio data (113), and the text data (115) provided to the data integration unit (190) may include time series information (173), and may be data in which values are listed according to the flow of time.
[0204] Excluding time information and utilizing sequence information may result in the loss of important information. Therefore, when integrating the video data (111), audio data (113), and text data (115) in the data integration unit (190), time series information must be taken into consideration.
[0205] In time series information (173), the time interval refers to the time difference between when information is stored. If time series information is integrated without considering the time interval, problems may arise. For example, if data with a 5-minute cycle and data with a 7-minute cycle are integrated, the cycle of the integrated data will be irregular, with intervals of "5, 7, 10, 14."
[0206] Additionally, in order to integrate different data in the data integration unit (190), it is necessary to check the data collection period. That is, if data collected from the 1st to the 5th and data collected from the 3rd to the 5th are integrated, data from the 1st to the 2nd days become missing values.
[0207] Figure 15 may be an example of generating integrated data (191) by integrating different data.
[0208] When different data 1 and data 2 are data collected from different devices over a period of one year from January 1, 2022 to December 31, 2022, the collection periods of said data 1 and said data 2 may be the same.
[0209] In addition, if the data 1 is collected based on 10 minutes and the data 2 is collected based on 5 minutes, the cycles of these two data do not match, so an integration cycle determination (193) for forming integrated data (191) can be performed. In addition, since the collection periods of the data 1 and the data 2 are the same, an integration cycle determination (195) may not be performed.
[0210] In this way, the integration cycle determination (193) can be determined as the greatest common divisor of the collection cycles of the data 1 and the data 2. In addition, the integration period determination (195) can be determined as a common section of the collection period of the data 1 and the data 2.
[0211] In this way, integrated data (191) can be formed by comparing the cycle and period in which the video data (111), the audio data (113), and the text data (115) are collected through the time series information (173) of each of the video data (111), the audio data (113), and the text data (115).
[0212] Referring to FIGS. 16 and 17, a data acquisition step (S100) may be included to generate multimodal data including photographing data (311), recording data (331), environmental data (351), and biometric data (371) through a photographing device (310), a recording device (330), an environmental sensor device (350), and a biometric sensor device (370) included in a data acquisition unit (300).
[0213] Through this, information in various formats can be acquired through various devices. In addition, a data processing step (S200) can be performed in which the photographing data (311), the recording data (331), the environmental data (351), and the biometric data (371) are integrated in the data processing unit (100) to generate integrated data (191).
[0214] Through this, each multimodal data acquired from various devices can be integrated into one data according to one time series information (173).
[0215] In addition, in the data study (500), a data processing step (S300) can be performed in which artificial intelligence analyzes the integrated data (191) to generate a result value required by the user.
[0216] Through this, various data can be analyzed by a single artificial intelligence, improving the consistency and accuracy of the analysis results, and the accuracy, diversity, and reliability of the results provided to the user can be improved as they are generated from various data.
[0217] Additionally, a provision step (S400) may be performed to provide the result values in a form that is easy for the user to check. In the provision step (S400), each result value is visualized, allowing the user to easily check the result values in various ways, including graphs, icons, and animations.
[0218] The above data processing step (S200) may proceed with a management step (S210) of classifying the photographing data (311), the recording data (331), the environmental data (351), and the biometric data (371) generated in the above data acquisition step (S100) by type and classifying them into image data (111), audio data (113), and text data (115).
[0219] Through this, it is possible to provide a processing process appropriate to the form for the photographing data (311), the recording data (331), the environmental data (351), and the biometric data (371), each having a different form.
[0220] A missing value confirmation step (S220) for determining missing values for each of the above image data (111), the above audio data (113), and the above text data (115) may be performed. In addition, a missing value processing step (S230) for processing each of the above image data (111), the above audio data (113), and the above text data (115) according to the missing values to generate processing data (143) may be performed.
[0221] Through this, by removing missing parts from the data acquired from each device, it is possible to prevent the artificial intelligence from being unable to derive results or deriving incorrect results due to missing values for the data acquired from each device.
[0222] By evaluating the quality of the above-mentioned processing data (143), an evaluation step (S240) can be performed to classify low-quality data (161) and usage data (163). Through this, data that can be analyzed by the artificial intelligence can be analyzed directly by the artificial intelligence, and improvements can be made to low-quality data (161) that cannot be analyzed.
[0223] An augmentation step (S250) can be performed to augment the above-mentioned low-quality data (161) to generate improved data (171) with improved quality. Through this, data with low quality that is difficult for artificial intelligence to analyze can be converted into data that can be analyzed by artificial intelligence.
[0224] An integration step (S260) can be performed to generate integrated data (191) by integrating each of the improvement data (171) and the usage data (163) generated from the above image data (111), the above audio data (113), and the above text data (115) into a time series.
[0225] By doing so, the load on the artificial intelligence provided in the data study (500) can be reduced by providing one integrated data to the data study (500), and the consistency of the result value can be secured by analyzing the data through one artificial intelligence.
[0226] In addition, the above integrated data (191) includes data acquired from various devices, thereby ensuring the diversity and reliability of the results analyzed by the artificial intelligence.
[0227]
[0228] While the present invention has been described with reference to the embodiments illustrated in the drawings, these are merely exemplary. Those skilled in the art will appreciate that various modifications may be made therefrom, and that all or part of the described embodiments may be selectively combined to form a configuration. Therefore, the true scope of technical protection of the present invention should be defined by the technical spirit of the appended claims.
[0229]
[0230] [Explanation of symbols]
[0231] 1: Data processing system
[0232] 100: Data Processing Department 110: Data Management Department
[0233] 111: Video data 113: Audio data
[0234] 115: Text data 130: Missing value check section
[0235] 131: Good data 133: Supplementary data
[0236] 135: Discarded data 140: Missing value processing section
[0237] 141: KNN Algorithm 143: Processing Data
[0238] 160: Data Evaluation Department 161: Low-quality data
[0239] 163: Usage Data 164: Text Inspection Method
[0240] 164a: Grammar error check 164b: Local consistency check
[0241] 164c: Global consistency check 164d: Readability check
[0242] 165: Spectogram 166: Brightness-excluded drawing
[0243] 167: HSV color model 167a: Color tone drawing
[0244] 167a: Saturation diagram 169: Blur diagram
[0245] 170: Data Augmentation Section 171: Improved Data
[0246] 173: Time series information 190: Data integration department
[0247] 191: Integrated Data 193: Determining the Integration Cycle
[0248] 195: Determining the integration period
Claims
1. Data acquisition unit including a photographing device, a recording device, an environmental sensor device, and a biosensor device; A data processing unit that receives each data acquired from the above data acquisition unit and creates integrated data; and Data study, which analyzes the integrated data through artificial intelligence to derive results; A system for integrating and augmenting multimodal data, characterized in that the data processing unit determines the quality of each data acquired from the data acquisition unit, improves the quality of each data, and generates the integrated data.
2. In paragraph 1, The above data processing unit, A data management unit that classifies multiple different data provided from the above data acquisition unit by type; A missing value verification unit that determines each missing value for multiple classified data; A missing value processing unit that processes and creates processing data according to the missing values of each data; A data evaluation department that evaluates the quality of the above processing data and determines whether it is low-quality data that is difficult to analyze through artificial intelligence and whether it is usable data that can be analyzed; A data augmentation unit that improves the quality of the above low-quality data to generate improved data; and A system for integrating and augmenting multimodal data, characterized by including a data integration unit that generates integrated data by integrating the improvement data and the usage data generated according to each data classified by type.
3. In paragraph 2, The above data management department, A system for integrating and augmenting multimodal data, characterized in that the data acquired from the above data acquisition unit are classified into image data, audio data, and text data according to their formats.
4. In paragraph 2, The above missing value confirmation section is, A system for integrating and augmenting multimodal data, characterized in that for each classified data, data with less than 10% missing values are classified as good data, data with 10% or more but less than 50% missing values are classified as supplementary data, and data with 50% or more missing values are classified as discarded data.
5. In paragraph 4, The above missing value processing unit is, A system for integrating and augmenting multimodal data, characterized in that the missing portions in the above-mentioned good data are deleted, the missing portions of the above-mentioned supplementary data are replaced with surrounding values through the KNN algorithm, and the above-mentioned discarded data are deleted entirely.
6. In paragraph 3, The above data evaluation department, A system for integrating and augmenting multimodal data, characterized in that it performs at least one of a grammatical error check, a local consistency check, a global consistency check, and a readability check on the above text data.
7. In paragraph 6, The above grammar error check is, A system for integrating and augmenting multimodal data, characterized in that the grammatical error inclusion rate is obtained through the following [Mathematical Formula 1]. [Mathematical Formula 1] Here, is the total number of documents, appeared in the document If the th word is an error, 1 is returned, otherwise 0.
8. In paragraph 6, The above local consistency test is, A system for integrating and augmenting multimodal data, characterized in that the relevance is determined for adjacent different sentences, and the relevance is confirmed through the following [Mathematical Formula 2]. [Mathematical formula 2] Here, is the number of total sentences, silver The second sentence and This is the relationship between the second sentence.
9. In paragraph 6, The above global consistency check is, A system for integrating and augmenting multimodal data, characterized in that the relationship between each sentence is analyzed through the following [Mathematical Formula 3] to determine whether all sentences constituting the entire text data are related to one topic. [Mathematical Formula 3] Here, is the number of total sentences, silver The second sentence and This is the relationship between the second sentence.
10. In clauses 8 and 9, The above interrelationships are, A system for integrating and augmenting multimodal data, characterized in that the occurrence of the same word in different sentences is determined using [Mathematical Formula 4], and the mutual relationship between different sentences is corrected using distance weights according to the positions of the same word appearing in different sentences. [Mathematical Formula 4] Here, It returns 1 if at least one identical word appears in another sentence as a relationship, otherwise it is counted as 0. is the distance between different sentences in which the above word appears.
11. In paragraph 6, The above readability test is, A system for integrating and augmenting multimodal data, characterized in that the length of sentences and the length of paragraphs included in the text data are measured through the average length of sentences included in the text data, and the length of sentences and the length of paragraphs included in the text data are evaluated through the following [Mathematical Formula 5]. [Mathematical Formula 5] Here, is the average paragraph length, is the average sentence length, and the larger the calculated value, the higher the level of the text data is considered to be.
12. In paragraph 3, The above audio data is, A system for integrating and augmenting multimodal data, characterized in that the signal-to-noise ratio included in the above audio data is converted into a spectrogram and evaluated by the data evaluation unit.
13. In paragraph 3, The above video data is, The data evaluation section above judges the quality of the image through the degree of image clarity and blurring. The above clarity is a system for integrating and augmenting multimodal data, characterized in that it determines the values of hue and saturation using the HSV color model.
14. In paragraph 2, The above data augmentation unit is, For the above low-quality data, if the low-quality data is in the form of text or image, the data is improved and generated through a deep neural network. A system for integrating and augmenting multimodal data, characterized in that if the above low-quality data is in audio form, the low-quality data in audio form is imaged, and improved and generated in the same way as the low-quality data in video form to generate the improved data.
15. In paragraph 2, The above data integration unit, The above utilization data and the above improvement data, which are formed by augmenting each data classified by type, further include time series information. A system for integrating and augmenting multimodal data, characterized in that each of the above data generates integrated data by matching the cycle and period in which each data is created based on time series information included in the above usage data and the above improvement data.
16. In paragraph 2, A server is formed to transmit external data to the data management unit to supplement each data measured in the data acquisition unit. A system for integrating and augmenting multimodal data, characterized in that the external data is classified according to the form in which the external data is composed in the data acquisition unit.
17. Since there is a second clause, Each data measured in the above data acquisition unit is stored to form a database that creates existing data. A system for integrating and augmenting multimodal data, characterized in that existing data measured in the past is transmitted to the data management unit and classified according to the form in which the existing data is composed.
18. In a system that integrates and augments multimodal data, A data acquisition step for generating photographing data, recording data, environmental data and biometric data through a photographing device, a recording device, an environmental sensor device and a biometric sensor device included in a data acquisition unit; A data processing step in which the data processing unit integrates the above-mentioned shooting data, the above-mentioned recording data, the above-mentioned environmental data, and the above-mentioned biometric data to generate integrated data; and A data processing method characterized by including a data processing step in which artificial intelligence analyzes the integrated data in the data study to derive result values required by the user.
19. In Article 18, The above data processing step is, A management step of classifying the photographing data, the recording data, the environmental data and the biometric data generated in the data acquisition step into image data, audio data and text data by type; A missing value confirmation step for determining missing values for each of the above image data, the above audio data, and the above text data; A missing data processing step for generating processing data by processing each of the image data, the audio data, and the text data according to the missing data; An evaluation step for evaluating the quality of the above processing data and classifying low-quality data and usage data; An augmentation step for augmenting the above low-quality data to generate improved data with improved quality; A data processing method characterized by including an integration step of generating integrated data by integrating each of the improvement data and the usage data generated from the image data, the audio data, and the text data into a time series.
Citation Information
Patent Citations
Multimodal Data Fusion Using Recurrent Neural Networks
JP2023501469A
Apparatus and method for processing multimodal fusion
KR1020080051479A
Apparatus and method for assessing data quality for text analysis
KR102019207B1
Automatic parts ordering device and method considering part failure prediction and aging rate determination
KR1020230005448A
Methods and systems for data management, integration, and interoperability
WO2022082095A1
Cited By
Method and system for intelligently generating chronic disease follow-up table based on multi-modal data fusion
CN120783927A
Medical image report automatic analysis method and system based on deep learning
CN122491245A