Multi-modal data governance method, device and equipment based on multi-agent cooperation

CN122173454BActive Publication Date: 2026-09-25SHANG HAI ZHANG JIANG SHU XUE YAN JIU YUAN
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610628092.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-09
Publication Date
2026-09-25
Estimated Expiration
2046-05-09

AI Technical Summary

Technical Problem

当面对海量且物理存储格式混乱的混合型数据时,传统计算机系统往往因无法自动解析数据类型而导致处理进程中断、内存溢出或统计算法加载失败

Benefits of technology

[0053]上述基于多智能体协作的多模态数据治理方法、装置、计算机设备、计算机可读存储介质和计算机程序产品,通过混合数据智能体对混合模态变量进行重定义与实体转换,在电数字数据处理阶段避免了因数据类型冲突引发的系统报错。同时,通过文本智能体建立映射关系并进行标准术语重写,大幅降低了离散变量的维度数,有效减少了后续统计算法在运行时对计算机内存的占用开销,提升了整体运算执行速度,提高了数据治理系统的运行稳定性与计算效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122173454B_ABST
    Figure CN122173454B_ABST
Patent Text Reader

Abstract

The application relates to a multi-modal data governance method, device and equipment based on multi-agent cooperation. The method comprises the following steps: identifying the variable mode of a variable in to-be-governed data; converting mixed-mode variables into text or numerical-mode variables by using a mixed data agent; performing standardization evaluation, standard term rewriting and material type division on the text-type variable by using a text agent; dividing the numerical-type variable into a continuous type, an unordered discrete type or an ordered discrete type based on variable semantics and value characteristics by using a numerical agent; and performing corresponding statistical algorithm governance operations according to the divided material types. By using the method, semantic-level recognition and conversion of multi-modal data can be realized through multi-agent cooperation, a governance strategy can be accurately matched based on fine statistical attribute classification, and the accuracy and automation efficiency of data governance are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic digital data processing technology, and in particular to a multimodal data governance method, apparatus, computer equipment, computer-readable storage medium, and computer program product based on multi-agent cooperation. Background Technology

[0002] With the deepening development of big data and artificial intelligence technologies in professional fields such as healthcare and finance, massive amounts of heterogeneous and multimodal data have emerged. While this type of data is easy to collect and can directly reflect actual business information, problems such as non-standardized descriptions and inconsistent recording and storage formats limit its application in subsequent modeling and analysis. Therefore, data governance techniques are needed to transform raw unstructured data into a standardized, machine-readable format.

[0003] Existing data governance solutions often focus on building up upper-level business rules, paying less attention to the computational overhead and stability issues of data governance systems when dealing with heterogeneous and multimodal data. When faced with massive amounts of mixed data with inconsistent physical storage formats, traditional computer systems often fail to automatically parse data types, leading to processing interruptions, memory overflows, or statistical algorithm loading failures.

[0004] Therefore, there is a need for a technical solution that can automatically detect features, dynamically switch modes, and optimize the physical storage structure of data at the level of electronic digital data processing, so as to improve the efficiency, robustness, and automation of computer systems in processing complex data. Summary of the Invention

[0005] Therefore, it is necessary to provide a method, apparatus, computer device, computer-readable storage medium, and computer program product for multimodal data governance based on multi-agent collaboration for hybrid data, addressing the aforementioned technical problems.

[0006] Firstly, this application provides a multimodal data governance method based on multi-agent cooperation, including:

[0007] Acquire the data to be governed, scan the initial variable values ​​of each variable in the data to be governed, and identify the variable modality of the variable as text modality, numerical modality or mixed modality based on the value distribution of the initial variable values;

[0008] The hybrid data agent is used to redefine and transform the hybrid variables belonging to the hybrid modality, converting them into text modal variables or numerical modal variables;

[0009] Obtain text-type variables whose variable modality is the text modality, wherein the text-type variables include variables whose variable modality is identified as text modality and the text modality variables obtained by conversion;

[0010] The text-based intelligent agent is used to perform standardized evaluation and corresponding processing on the text-based variables, and the data types of the text-based variables are classified into non-standard text type, unordered discrete type, or ordered discrete type.

[0011] Obtain numerical variables whose variable mode is the numerical mode, wherein the numerical variables include variables whose variable mode is identified as numerical modes and the numerical mode variables obtained by conversion;

[0012] Using numerical intelligent agents, the data types of the numerical variables are classified into continuous, unordered discrete, or ordered discrete types based on the numerical variable names and values.

[0013] Based on the data type classification, the corresponding statistical algorithm is executed to perform the data management operation and generate the managed data.

[0014] In one embodiment, the step of scanning the initial variable values ​​of each variable in the data to be managed, and identifying the variable modality of the variable as text modality, numerical modality, or mixed modality based on the value distribution of the initial variable values, includes:

[0015] Traverse the initial variable values ​​of each variable in the data to be managed to identify character distribution characteristics;

[0016] Based on the character distribution characteristics, determine the variable mode of the variable:

[0017] If the initial variable values ​​are all composed of numeric characters, the variable mode of the corresponding variable is determined to be a numeric mode.

[0018] If the initial variable values ​​are all composed of non-numeric characters, the variable modality of the corresponding variable is determined to be text modality;

[0019] If the initial variable value contains both numeric characters and non-numeric characters, the variable mode of the corresponding variable is determined to be a mixed mode.

[0020] In one embodiment, the redefinition and transformation of hybrid variables belonging to the hybrid modality using a hybrid data agent includes:

[0021] Read the mixed variable name and mixed variable value of the mixed type variable;

[0022] Identify the semantic meaning of the mixed variable names and the distribution characteristics of the numerical content in the mixed variable values;

[0023] Based on the semantic orientation and the distribution characteristics, the hybrid variable is redefined as a text modal variable or a numerical modal variable;

[0024] When redefined as a text modal variable, the original content of the mixed variable value is retained as the text variable value of the transformed text modal variable;

[0025] When redefined as a numerical modal variable, the numerical content of the mixed variable value is extracted as the numerical variable value of the transformed numerical modal variable.

[0026] In one embodiment, the standardization evaluation and corresponding processing of the text-type variables using a text agent includes:

[0027] Analyze the distribution patterns of the text variable values ​​for all the aforementioned text-type variables;

[0028] If the distribution pattern is not statistically regular, the standardized evaluation result is determined to be unnecessary for standardization, and the data type of the text variable is classified as non-standard text type.

[0029] If the distribution pattern exhibits statistical regularity, the standardized evaluation result is determined to require standardization. Based on the textual variables, a standard terminology mapping is generated, and the textual variable values ​​are rewritten. The rewritten textual variables are then classified into unordered discrete or ordered discrete data types, where:

[0030] The process of generating standard terminology mappings and rewriting text variable values ​​includes performing semantic clustering on the text variable values, establishing a mapping relationship between heterogeneous words and standard terms, and replacing the original text variable values ​​with the standard terms.

[0031] The classification of data types for the rewritten text variables includes:

[0032] Analyze whether there is a semantic-based hierarchical order among the standard terms;

[0033] If the specified hierarchical order exists, the data type of the text variable is classified as ordered discrete.

[0034] If the hierarchical order is zero, the data type of the text variable is classified as unordered discrete.

[0035] In one embodiment, the step of using a numerical agent to classify the data type of the numerical variables into continuous, unordered discrete, or ordered discrete types based on the numerical variable names and values ​​includes:

[0036] Analyze the semantics of the numerical variable names, and statistically analyze the value characteristics and rank order of the numerical variable values;

[0037] In the case where the semantics of the indicator are continuously changing, the value characteristics are continuous interval values, and the level order is nonexistent, the numerical variable is classified as continuous.

[0038] In the semantic-oriented classification index, where the value feature is discrete numerical and the rank order is nonexistent, the numerical variable is classified as unordered discrete.

[0039] In the semantic-oriented hierarchical index, where the value feature is discrete numerical and the hierarchical order is present, the numerical variable is classified as an ordered discrete type.

[0040] In one embodiment, the step of performing corresponding statistical algorithm governance operations based on the data types to generate governance data includes:

[0041] Outlier identification, missing value imputation, and standardization transformation are performed on the continuous variables, where the continuous variables are numerical variables whose data type is continuous.

[0042] Missing value imputation and small sample class merging operations are performed on the unordered discrete variables. The unordered discrete variables include text variables and numerical variables whose data type is unordered discrete.

[0043] Missing value imputation and small sample adjacent category merging operations are performed on the ordered discrete variables. The ordered discrete variables include text variables of the data type ordered discrete and numerical variables of the data type ordered discrete.

[0044] Secondly, this application also provides a multimodal data governance device based on multi-agent cooperation, comprising:

[0045] The data exploration module is used to acquire the data to be managed, scan the initial variable values ​​of each variable in the data to be managed, and identify the variable modality of the variable as text modality, numerical modality or mixed modality according to the value distribution of the initial variable values.

[0046] The hybrid processing module is used to redefine and transform hybrid variables belonging to the hybrid modality using a hybrid data agent, converting them into text modality variables or numerical modality variables;

[0047] The text processing module is used to obtain text-type variables whose variable modality is the text modality. The text-type variables include variables whose variable modality is identified as text modality and the text modality variables obtained by conversion. The text-type variables are standardized and processed using a text agent to classify the data type of the text-type variables into non-standard text type, unordered discrete type, or ordered discrete type.

[0048] The numerical processing module is used to obtain numerical variables whose variable mode is the numerical mode, the numerical variables including variables whose variable mode is identified as numerical mode and the numerical mode variables obtained by conversion; and to use a numerical agent to classify the data type of the numerical variables into continuous, unordered discrete or ordered discrete types based on the numerical variable name and numerical variable value of the numerical variables.

[0049] The data cleaning module is used to perform corresponding statistical algorithm processing operations based on the data types to generate processed data.

[0050] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.

[0051] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.

[0052] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.

[0053] The aforementioned multimodal data governance methods, devices, computer equipment, computer-readable storage media, and computer program products based on multi-agent collaboration redefine and transform hybrid modal variables through hybrid data agents, avoiding system errors caused by data type conflicts during the electronic digital data processing stage. Simultaneously, by establishing mapping relationships and rewriting standard terminology through text agents, the dimensionality of discrete variables is significantly reduced, effectively minimizing the memory overhead of subsequent statistical algorithms, improving overall computational speed, and enhancing the operational stability and computational efficiency of the data governance system.

[0054] Secondly, by leveraging text-based and numerical agents to deeply understand the semantics of variable names and extract the value characteristics of the data, the original variables are mapped to continuous, unordered discrete, or ordered discrete variables with statistically defined types. Further type classification based on data characteristics enables the processed dataset to better adapt to corresponding statistical algorithms. This avoids the waste of processor resources caused by type misjudgment in traditional methods, such as performing floating-point averaging operations on unordered discrete gender codes, thus optimizing machine readability and adaptability to automated statistical algorithms. Attached Figure Description

[0055] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0056] Figure 1 This is a flowchart illustrating a multimodal data governance method based on multi-agent collaboration in one embodiment;

[0057] Figure 2 This is a flowchart illustrating the redefinition and transformation of mixed-modal variables in one embodiment;

[0058] Figure 3 This is a flowchart illustrating the processing of text variables in one embodiment;

[0059] Figure 4 This is a flowchart illustrating the processing of numerical variables in one embodiment;

[0060] Figure 5 This is a structural block diagram of a multimodal data governance device based on multi-agent collaboration in one embodiment;

[0061] Figure 6 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0062] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0063] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0064] Existing data governance techniques have significant technical shortcomings when dealing with mixed data that intertwines text and numerical values. First, existing algorithms lack an adaptive splitting mechanism based on the inherent characteristics of the data. In real-world datasets, mixed data often exhibits different semantic attributes; for example, some variables are essentially numerical but contain textual descriptions, while others are essentially textual but contain numbers. Existing techniques typically employ a single strategy of forcibly converting data to numerical or textual values, failing to automatically identify the true attributes of variables. This leads to the loss of key semantics or incorrect type labeling in the extracted data, consequently preventing subsequent statistical algorithms from correctly loading the data.

[0065] Secondly, existing automated data governance systems lack flexible rule injection and parameter correction interfaces. Fully automated models typically operate based on closed algorithmic logic. When faced with specific data or edge cases with low algorithm confidence, they cannot dynamically retrieve externally pre-defined imputation rules, nor can they respond to external correction instructions to update erroneous data labels. This closed system architecture makes it difficult to eliminate accumulated errors during the data governance process, severely impacting the usability and accuracy of the final dataset.

[0066] For dirty data in structured tables—that is, fields that should be numerical or short text but contain mixed content—existing technologies typically either extract all numbers using regular expressions or treat all data as text. This application addresses the accuracy and logical consistency issues in the automated cleaning and governance of multimodal data in existing technologies by proposing a data governance method that dynamically determines whether to purify numbers or retain the original values ​​based on variable name semantics and value distribution characteristics. For example, a statement like "once a day" is classified as text and its original value is retained based on the variable name semantics of "frequency" and the value distribution characteristic of "text modality."

[0067] The data governance method provided in this application can be applied to any electronic device with data processing capabilities. Electronic devices include, but are not limited to:

[0068] Terminal devices, such as desktop computers, laptops, workstations, and tablets, are suitable for individual researchers or small-scale data governance scenarios. Users can directly import comma-separated values ​​(CSV) files locally for data governance. Data acquisition, agent inference, and statistical algorithm execution are all completed on the same local device, ensuring that data does not leave the domain and meeting high privacy requirements, such as patient privacy data.

[0069] Server-side equipment, including standalone physical servers, server clusters, distributed systems, or cloud servers, is suitable for large-scale data governance centers in medical institutions, research institutes, or enterprises, capable of handling massive concurrent electronic medical records or business data. The client side is responsible for data collection and result display, while the server side is responsible for running intelligent agents and data cleaning algorithms.

[0070] The cloud platform, a cloud computing platform based on virtualization technology, provides data governance services in the form of software as a service or platform as a service. Users upload data to be governed through network interfaces, and the cloud cluster schedules multiple intelligent agent instances to process the data columns in parallel and returns the governed dataset to the front end.

[0071] Electronic devices typically include a processor, memory, and a communication interface. The processor executes computer program instructions stored in the memory to implement the multi-agent collaborative governance logic of this application; in particular, for semantic understanding tasks involving numerical agents, text agents, and hybrid data agents, a graphics processing unit (GPU) or tensor processor can be used to accelerate inference of large language models.

[0072] The intelligent agent mentioned in this application, including hybrid data intelligent agents, text intelligent agents, and numerical intelligent agents, refers to a software module or logical unit with perception, decision-making, and execution capabilities. Its implementation includes: driven by a large language model, the intelligent agent understands the semantics of variable names and the intent of variable values ​​by calling the application programming interface of a pre-trained large-scale language model or a locally deployed model, such as a transformer architecture model, a transformer-based bidirectional encoder representation model, a series of generative pre-trained models, a series of general language models, or a vertical domain model specific to the medical field, using cue word engineering. Further, it can be combined with knowledge graphs, by referring to industry knowledge graphs or standard terminology sets, such as the International Statistical Classification of Diseases and Health-Related Issues, Medical Systems Nomenclature - Clinical Terminology, or Medical Subject Headings, to perform standardized terminology mapping. For numerical intelligent agents, in addition to semantic understanding, statistical rule patterns or lightweight machine learning classifiers can also be combined to determine continuous or discrete types.

[0073] The data to be processed in this application can originate from various channels and formats. Sources may include electronic health record systems, hospital information systems, laboratory information management systems, clinical trial data acquisition systems, enterprise resource planning systems, customer relationship management systems, or IoT sensor logs. Data formats may include structured data files, database exported files, or unstructured text data streams after preliminary parsing. Data content may include any tabular data containing multimodal characteristics, such as clinical diagnosis and treatment data, demographic data, financial transaction logs, or logistics supply chain data, encompassing numerical values, text, and a mixture of both.

[0074] In one exemplary embodiment, such as Figure 1 As shown, a multimodal data governance method based on multi-agent cooperation is provided. Taking the application of this method to a single-machine architecture as an example, the method includes the following steps S110 to S150. Wherein:

[0075] Step S110: Obtain the data to be governed, scan the initial variable values ​​of each variable in the data to be governed, and identify the variable modality of the variable as text modality, numerical modality or mixed modality according to the distribution of the initial variable values.

[0076] Among them, data to be cleaned refers to raw datasets that need to be cleaned, transformed, or standardized. It can be in the form of spreadsheet files stored on storage media, database query result sets, comma- or tab-separated text files, or data streams received through data interfaces.

[0077] The data to be processed can be medical and health records, industrial production logs, financial transaction records, or any structured or semi-structured data containing multi-dimensional attribute information. It usually presents a quasi-structured or structured form, but there are problems such as messy data annotation information, synonyms, inconsistent missing information, information hidden in long strings of text, redundant information, or abnormal outliers that do not conform to common sense.

[0078] Variable modality is a classification description of the physical representation of data. Numerical modality refers to data consisting of Arabic numerals, textual modality refers to data consisting of Chinese characters, English letters or other characters, and mixed modality refers to data items in the same variable that contain both numerical and textual data, or a single data item that contains both numerical and non-numerical characters.

[0079] While acquiring data, it is also possible to obtain auxiliary descriptive information associated with the data, such as a variable dictionary. This variable dictionary can also be called a metadata definition table, data schema file, or field description document. Its purpose is to provide a mapping between variable names and their actual business meanings. Especially when variable naming rules are meaningless, such as "var1", the true semantics of the variable can still be interpreted based on its corresponding variable meaning "family history".

[0080] Identifying variable modalities refers to the preliminary classification of the physical attributes of a data column. This identification can be achieved in various ways, such as scanning based on regular features to statistically analyze the distribution ratio of numeric characters, alphanumeric characters, and special symbols in the column data and determining the modality based on the ratio threshold; or inferring based on the type of the data interface by directly reading the storage type of the column using the data analysis library built into the programming language; or using sampling probing to read the first few rows of samples in the data column for type prediction.

[0081] Step S120: Use the hybrid data agent to redefine and transform the hybrid variables belonging to the hybrid modality, and convert them into text modal variables or numerical modal variables.

[0082] Hybrid data agents refer to logical units used to process data with complex formats. These can be pre-trained artificial intelligence models, parsing units built based on preset logical rules, or regular expression processors. Redefinition and transformation processing refers to converting structurally ambiguous hybrid data into structurally defined single-type data.

[0083] In practical implementation, the agent can use a labeling method to generate a new type label for the variable and distribute the original data to different processing queues according to the label; it can also use an extraction method to separate the numerical part from the mixed string as a new numerical modal variable, and separate the unit or note information as a new text modal variable; it can also use a mapping method to convert non-numerical characters into specific numerical codes or null placeholders according to preset rules, thereby uniformly converting them into numerical modalities.

[0084] Through the above processing, the originally ambiguous mixed-modal variables are transformed into text-modal or numerical-modal variables that can be processed in subsequent steps. For mixed-modal data that traditional data governance tools (Extract-Transform-Load, or ETL) struggle to handle—that is, data where the same field contains both numerical and textual data—this embodiment introduces a mixed-data agent for redefinition and transformation, rather than simply and crudely eliminating or forcibly converting it. This allows complex data that would otherwise be discarded or flagged as errors to be transformed into usable text and numerical variables, significantly preserving the business value of the original data and improving the information completeness of the governed dataset.

[0085] Step S130: Obtain text-type variables whose variable modality is text mode, and use a text agent to perform standardized evaluation and corresponding processing on the text-type variables, classifying the data type of the text-type variables into non-standard text type, unordered discrete type, or ordered discrete type.

[0086] The text-type variables include variables identified as text modalities and the transformed text modal variables. It should be understood that the text-type variables in this step include both the original variables directly identified as text modalities in step S110 and the new text modal variables generated after transformation in step S120. The text agent is a module with semantic understanding capabilities, which can be implemented by calling knowledge graphs, thesaurus, or semantic vector models.

[0087] Specifically, once the variable modality of the data to be governed is determined to be a text modality, the text agent first performs a standardization evaluation on this data. For disorganized descriptive text lacking statistical regularity, it is directly labeled as "non-standard text," and the evaluation result indicates that standardization is not required; that is, this type of data has no statistical analysis value and will not participate in subsequent in-depth governance processes. For text variables that require standardization, the process of generating standard terminology mappings and rewriting variable values ​​can also be called data normalization, terminology standardization, or data cleaning. This can be achieved by identifying synonyms, correcting spelling errors, standardizing abbreviations, or converting colloquialisms into standard professional terms.

[0088] Furthermore, classifying variables into unordered or ordered discrete types refers to determining their statistical type based on their semantic attributes. Unordered discrete variables, also known as nominal or categorical variables, indicate that there is no inherent hierarchical or ordinal relationship between variable values. Ordered discrete variables, also known as ordinal or ranked variables, indicate that there is a logical relationship of magnitude, degree, or sequence between variable values. This classification can be achieved by analyzing the semantics of variable values, such as high, medium, and low, or by consulting the definitions in a variable dictionary.

[0089] Step S140: Obtain numerical variables whose variable mode is numerical mode, and use the numerical agent to classify the data type of numerical variables into continuous, unordered discrete, or ordered discrete types based on the numerical variable name and numerical variable value.

[0090] Numerical variables include variables whose variable modes are identified as numerical modes and the numerical mode variables obtained through transformation.

[0091] Similarly, the numerical variables in this step encompass both the original numerical modal variables identified in step S110 and the numerical modal variables transformed in step S120. A numerical agent is a logical module used to analyze numerical features, which can be implemented using statistical analysis algorithms or classification models.

[0092] For example, numerical intelligent agents classify variables based on the business meaning implied by the names of numerical variables and the distribution characteristics of the values ​​of numerical variables. The continuous type after classification can also be called quantitative variable or scalar, which means that the variable can take any value within a certain range and usually has mathematical operation meaning. The unordered discrete type here specifically refers to variables that are composed of numbers but are actually only used as identifiers or codes, such as "gender code" or "whether or not to smoke". The ordered discrete type here specifically refers to variables that use numbers to represent levels or scores, such as satisfaction scores of 1 to 5.

[0093] Step S150: Perform the corresponding statistical algorithm processing operation according to the data type to generate processed data.

[0094] Among them, statistical algorithm governance refers to the process of optimizing data quality by applying mathematical or statistical methods. Based on the sub-types determined in the aforementioned steps, different governance strategies are automatically matched to variables defined as continuous, unordered discrete, and ordered discrete.

[0095] These governance strategies include, but are not limited to: handling outliers with anomalous distributions, such as truncation, removal, or smoothing; imputation of missing values, such as interpolation, estimation, or filling in specific values; and merging small samples of data categories, such as classification, dimensionality reduction, or reorganization.

[0096] The governance dataset generated after the governance data is collected is a cleaned, standardized and structured collection of data. It can be directly used for subsequent data mining, machine learning modeling or statistical analysis tasks, or it can be stored in a data warehouse as an asset.

[0097] The aforementioned data governance method first uses a hybrid data agent to convert mixed modal variables (where numbers and text are mixed) into text modal variables or numerical modal variables, preserving all useful information. By combining the semantic connotations of variable names (e.g., "creatinine" implies numerical values, and "remarks" implies text), the mixed modality is accurately identified and redefined. This method can automatically and losslessly convert complex data, such as numerical values ​​with units or records containing numerical values, into independent numerical or text variables. It avoids the problem of entire columns being removed or forcibly emptyed due to the inability to parse the data, as in traditional methods. While ensuring smooth data processing, it maximizes the preservation of the original data's business value.

[0098] Furthermore, the aforementioned data governance methods utilize text-based and numerical agents to understand the meaning of variable names and content. This not only distinguishes between numbers and text but also further categorizes data types. It can accurately identify data that is "numerical but lacks numerical significance," such as gender codes, or "textual but with hierarchical order," such as high, medium, and low. This semantic-based classification corrects misclassifications caused by relying solely on storage format classification; for example, it identifies gender codes as continuous numerical values, providing a reliable reference for subsequent automated matching and data cleaning algorithms.

[0099] The aforementioned data governance method combines terminology standardization, type inference, and statistical cleaning algorithms. The intelligent agent can automatically generate standard terminology mappings and rewrite variable values, reducing the tedious process of manually writing cleaning rules. The execution of adaptive statistical algorithms, such as automatically merging small samples and automatically filling in missing values, further improves the efficiency and consistency of massive data governance.

[0100] In an exemplary embodiment, taking the governance of multimodal electronic medical record data in the field of clinical medicine as an example, the data governance method of this application is described in detail, aiming to further illustrate the technical solution of this application, rather than limiting the scope of protection of this application. In this embodiment, the multimodal data to be governed is a cervical lesion screening dataset from the gynecology department of a hospital, which includes variables of text modality, numerical modality and mixed modality, as shown in Table 1, the multimodal dataset to be governed. The variable names include identifier (ID), liquid-based thin-layer cytology test (TCT) and human papillomavirus (HPV), and the variable values ​​include atypical squamous cells not excluding high-grade lesions (ASC-H) and high-grade squamous intraepithelial lesions (HSIL).

[0101] Table 1

[0102]

[0103] In some embodiments, step S110 acquires the data to be governed, scans the initial variable values ​​of each variable in the data to be governed, and identifies the variable modality of the variable as text modality, numerical modality, or mixed modality based on the distribution of the initial variable values, including the following steps S112 to S114, wherein:

[0104] Step S112: Traverse the initial variable values ​​of each variable in the data to be processed and identify character distribution characteristics.

[0105] The initial variable values ​​are analyzed at the character encoding level. By scanning the encoding format or type of each character, the character distribution characteristics of the initial variable values ​​are identified. The character distribution characteristics are used to characterize whether the variable values ​​contain numbers, letters, symbols, or other encoded forms.

[0106] For example, a regular expression algorithm is used to traverse or sample the initial variable values ​​of each variable in the data to be processed, calculate the number of rows with successful matches and the number of rows with unsuccessful matches, and identify the distribution characteristics of numeric and non-numeric characters in the variable values; for example, scanning the variable value of "HPV result" identifies that it contains numeric characters such as "0" and "1", as well as non-numeric characters such as "no results found", "negative", and "positive", resulting in a distribution feature vector, such as [HPV result numeric character percentage: 92%, non-numeric character percentage: 5%, mixed character percentage: 3%]. The processing module outputs the character distribution characteristics.

[0107] Step S114: Determine the variable modality based on the character distribution characteristics. This determination process may include the following three parallel processing branches:

[0108] In the first processing branch, if the initial variable values ​​consist entirely of numeric characters, the variable mode of the corresponding variable is determined to be a numeric mode. This variable only contains quantized information and has no non-numeric interference terms. In this case, the variable is labeled as numeric so that it can be processed by the subsequent numerical agent.

[0109] In the second processing branch, if the initial variable values ​​consist entirely of non-numeric characters, the variable modality is determined to be text modality. This variable contains only descriptive information or non-quantitative symbols. In this case, the variable is labeled as text for subsequent text-based intelligent agent processing.

[0110] In processing branch three, if the initial variable value contains both numeric and non-numeric characters, the variable modality of the corresponding variable is determined to be a mixed modality. For example, for numeric values ​​containing units of measurement or text containing numerical codes, the variable modality of the corresponding variable is determined to be a mixed modality. In this case, the variable will be labeled as mixed-type so that the mixed data agent can access and process it later.

[0111] For example, the variable value scan for "HPV Result" revealed values ​​such as "No results found," "0 negative," "1 positive," and "0." Since the initial variable values ​​contained numeric characters, non-numeric characters, and mixed characters, it was automatically classified as a mixed modal variable. Combining the variable value scan results in Table 1, "TCT Result" and "Location" were ultimately determined to be text modal variables, "Age" to be a numeric modal variable, and "Marital Status," "HPV Result," and "Viral Load" to be mixed modal variables.

[0112] This embodiment, based on the actual value distribution of initial variables in the dataset to be governed, avoids the problem of discrepancies between manually predefined definitions and actual data. For example, it addresses situations where a variable is designed to be numerical but its initial values ​​contain a large amount of text. This ensures that the classification results reflect the true attributes of the data and avoids errors caused by modality misjudgment in subsequent processing. Furthermore, by classifying variables into modalities, corresponding processing modules can be directly matched, and corresponding processing measures can be taken, avoiding process chaos. If the initial values ​​of variables are all pure numerical, they are directly determined to be numerical modalities, eliminating the need for redundant calculations such as subsequent semantic analysis and format validation. If the initial values ​​are mainly in mixed formats, the conversion process is triggered first, reducing invalid processing steps, which is particularly suitable for the efficient governance needs of large-sample, multi-source data.

[0113] In some embodiments, after obtaining the data to be governed, the process further includes a preliminary exploration of the data to be governed, including:

[0114] First, keyword matching is performed on the column names to detect commonly used fields in the medical field and automatically identify the table headers. Variables with specific formats are identified based on the value characteristics of the data; for example, if all data in a column is in the format "YYYY-MM-DD", and the variable name semantics are combined to represent "date" or "time", this variable is determined as basic information. Additionally, variables such as place of origin and occupation are also considered basic information and are not involved in subsequent processing. Then, regular expressions are used to identify various common missing labels, such as "not seen", "NA", "None", "Nan", and "??", which, along with synchronously detected null values, are uniformly filled with the standard semantic representation "NA". Furthermore, if inherent filling rules are found in the variable dictionary, these rules are used to replace or fill the data.

[0115] Secondly, named entity recognition algorithms are used to identify potential personal privacy information in the dataset, such as names, ID cards, and residential addresses, and this sensitive information is deleted to complete the data privacy removal process. In addition, to ensure compatibility with repeated measurement data, when deleting personal privacy information, unique identification codes that can distinguish individuals, such as medical record IDs, are retained simultaneously, or unique IDs are pre-generated for differentiation.

[0116] Finally, by traversing the global data, rows or columns containing only "NA" cells are deleted; through hash comparison, duplicate rows or columns with identical values ​​are identified, retaining the first occurrence and deleting the rest; columns with constant variable values ​​are deleted, such as the gender column in survey data of pregnant women in late pregnancy where all values ​​are "female". After this preliminary data investigation, the data to be processed is initially cleaned, improving data utilization and processing efficiency in subsequent processing stages.

[0117] In some embodiments, to further improve the accuracy of governance, a variable dictionary associated with the data to be governed is pre-loaded before initiating multi-agent collaborative governance. The variable dictionary, as a priori knowledge base, stores variable names, descriptions of variable meanings, and inherent substitution rules for specific variables.

[0118] This embodiment provides a variable dictionary associated with Table 1; this variable dictionary contains variable names, variable meanings, and inherent substitution rules. See Table 2 for details.

[0119] Table 2

[0120]

[0121] Perform rule replacement or imputation based on the variable dictionary on all variables to be governed, specifically including:

[0122] First, the variable name of the variable being processed is parsed, and a match is searched in the variable dictionary. If a matching variable dictionary entry is found, the agent reads the inherent substitution rule under that entry. The inherent substitution rule defines a deterministic mapping relationship between the original data values ​​and the normalized values.

[0123] Secondly, a rule-first matching strategy is implemented. The agent iterates through the initial variable values, and for values ​​that exist in the inherent replacement rules, it directly replaces or fills them with normalized values ​​according to the rules, skipping the model inference process, thereby reducing computational overhead and ensuring the absolute accuracy of core terms.

[0124] For example, in the electronic medical record data compiled by a hospital department over the years, the variables of smoking, medical history, and family history only record those with positive values. Those with negative values ​​or random / non-random missing values ​​are not recorded. In this case, an inherent filling rule can be manually added to fill the missing values ​​of these variables with "no" or "0", indicating that the person does not smoke, has no medical history, or no family history, to ensure that the variable is correctly processed in the future.

[0125] Finally, multi-agent collaborative governance is initiated, proceeding to the subsequent governance process. By introducing a variable dictionary, this application can combine hard rule constraints with soft semantic understanding, ensuring governance flexibility while minimizing the risk of model illusion.

[0126] In some embodiments, such as Figure 2 As shown, step S120 utilizes a hybrid data agent to redefine and transform the hybrid modal variables, including steps S210 to S230, wherein:

[0127] Step S210: Read the mixed variable name and mixed variable value of the mixed type variable.

[0128] For example, the data reading program of the hybrid data agent is started to locate and read the hybrid variable name and the corresponding hybrid variable value from the dataset to be governed. The hybrid variable value is usually a combination of numerical and text characters, such as "once a day" and "0 negative".

[0129] Step S220: Identify the semantic meaning of the mixed variable name and the distribution characteristics of the numerical content in the mixed variable value.

[0130] For example, semantic analysis and feature extraction can be performed in parallel using the natural language processing module built into the hybrid data agent.

[0131] On the one hand, semantic parsing is performed on mixed variable names to identify their semantic orientation. Semantic orientation is used to determine whether the variable tends to describe quantitative attributes or qualitative identifier attributes.

[0132] On the other hand, the agent scans the mixed variable values, counts the frequency, positional patterns and effective proportion of numerical characters, calculates the information entropy weight of numerical content in the sample data, and analyzes the role of non-numerical characters in the data, such as units of measurement, state description suffixes or meaningless noise, thereby identifying the distribution characteristics of numerical content in the mixed variable values.

[0133] Step S230: Based on semantic orientation and distribution characteristics, redefine the hybrid variable as a text modal variable or a numerical modal variable.

[0134] The core information of the mixed variable is determined to be a text description or classification identifier when one of the following conditions is met: the semantic reference is clearly an identifier attribute; the ratio of the number of unique values ​​in the mixed variable to the total number of records is extremely low, and the numerical part has no computational meaning; the average proportion of the number of numeric characters in the mixed variable to the total number of characters is lower than a preset noise threshold.

[0135] For example, the variable "frequency of medication" can be determined by the semantics of the name and the characteristics of the data values ​​("once a day", "three times a week"), and only a complete text description can represent the core meaning of the value. Extracting only the numbers will lose key information, so the variable will be redefined as text data, and the mixed value will be retained as the original value.

[0136] The core information of a mixed variable is determined to be numerical when one of the following conditions is met: the semantics are clearly quantifiable; the mixed variable values ​​show high pattern consistency and the extracted numerical part has statistically significant distribution characteristics.

[0137] For example, the "HPV result" variable has mostly numerical values, and these values ​​represent the classification results; therefore, the semantic tendency is determined to be numerically dominant. Similarly, the "viral load" variable has mostly numerical values, and the remaining values ​​after removing the text can represent the specific viral load value; therefore, the semantic tendency is determined to be numerically dominant.

[0138] When redefined as a text modal variable, the original content of the mixed variable value is retained as the text variable value of the transformed text modal variable.

[0139] After generating the new decision type as a text variable, the text preservation instruction of the hybrid data agent is triggered. The decision value is only a part of the text identifier. The agent does not perform character stripping, but directly retains the original content of the hybrid variable value, mapping it completely to the text variable value of the transformed text modality variable.

[0140] When redefined as a numerical modal variable, the numerical content of the mixed variable values ​​is extracted as the numerical variable value of the transformed numerical modal variable.

[0141] After generating the new judgment type as a numerical variable, a numerical extraction instruction is triggered for the hybrid data agent. This instruction drives the agent to perform a cleaning operation on the original variable values. Taking "HPV result" which includes a state description as an example, when the original value is "0 negative" or "1 positive", the agent identifies "0" and "1" as core quantitative indicators, while "negative" and "positive" are redundant text. Finally, the agent performs a stripping operation, removing redundant text, retaining the core quantitative indicators, and converting the cleaned data into standard floating-point or integer formats, thus completing the entity conversion from hybrid modality to numerical modality. That is, the variable value of "HPV result" is converted to "NA" "0" "1" "1" "0" "0" "1", and the variable value of "viral load" is converted to "NA" "100" "1160" "298000" "10" "80" "200", and the variable modality of both variables is updated to numerical modality.

[0142] This embodiment utilizes a hybrid data agent to perform semantic tendency analysis on hybrid modal variables and selects a transformation strategy based on this analysis—either extracting numerical values ​​or preserving the whole—effectively solving the problem of mixed numerical values ​​and units in dirty data. Through this processing, variables that were originally unable to participate in mathematical operations due to the presence of non-numerical characters are cleaned into pure numerical modal variables, or ambiguous mixed content is standardized into textual descriptions. This process effectively preserves complex data that might be directly discarded in traditional governance processes, improving data utilization and information integrity.

[0143] In some embodiments, such as Figure 3As shown, step S130 utilizes a text agent to perform standardized evaluation and corresponding processing on text-type variables, including the following steps S310 to S330, wherein:

[0144] Step S310: Distribution pattern analysis and standardized evaluation.

[0145] Analyze the distribution patterns of text variable values ​​for all text-type variables. The text agent iterates through all text variable values, calculating the repetition rate and dispersion of the variable values. The agent then calculates the ratio of the number of different values ​​to the total number of records, as well as the frequency of each value.

[0146] If the distribution pattern is not statistically regular, the standardized evaluation result is determined to be unnecessary, and the data type of text variables is classified as non-standard text type.

[0147] Specifically, if the number of different values ​​is close to the total number of records, i.e., the ratio approaches 1, and the text content is mostly long text or natural language descriptions, such as "self-reported illness" or "past medical history," then the distribution pattern is determined to be statistically irregular. In this case, the evaluation results do not require standardization. The data type of this text variable is directly marked as non-standard text, and subsequent governance strategies can choose to ignore it or only perform basic cleaning.

[0148] If the distribution pattern shows statistical regularity, the standardized evaluation result is determined to require standardization.

[0149] Specifically, if the number of different values ​​is much smaller than the total number of records or the ratio is lower than a preset threshold, and the text content exhibits obvious enumeration or clustering characteristics, such as "male / female" or "negative / positive," the distribution pattern is determined to be statistically regular. In this case, a standardized evaluation result is generated, triggering the subsequent terminology rewriting process.

[0150] Step S320, standard terminology mapping and rewriting, specifically includes steps S322 to S326, wherein:

[0151] Step S322: Perform semantic clustering on the text variable values.

[0152] Natural language processing algorithms are used to group text variable values ​​with similar semantics but different expressions into the same cluster. In this embodiment, text-type variables include the variables "TCT result" and "location" whose variable modalities are identified as text modalities. The text agent embeds all the deduplicated values ​​in this column into vectors and projects them into a high-dimensional semantic space. The cosine similarity between vectors is calculated. For example, after clustering the variable values ​​of "TCT result", "few low-grade squamous intraepithelial lesions (LSIL)" and "low-grade squamous intraepithelial lesions (LSIL)" are identified as heterogeneous terms with the same semantics, and "no intraepithelial lesions or malignant cells" and "no malignant cells (NILM)" are also heterogeneous terms with the same semantics.

[0153] Step S324: Establish the mapping relationship between heterogeneous words and standard terms.

[0154] For each cluster, the agent identifies high-frequency words or determines a representative standard term based on a pre-defined industry standard dictionary. Subsequently, the agent establishes a mapping table between heterogeneous vocabulary and the standard term. It should be understood that this mapping can be generated by the agent searching common medical fields and words and comparing the tags in the data with those in a medical vocabulary database, or it can be pre-defined by someone skilled in the art.

[0155] In this embodiment, the agent queries the built-in standard terminology library based on the semantic context of the variable name "TCT result". For the "TCT result" variable, the determined standard terms are "NILM", "ASC-H", "LSIL", "HSIL", and "SCC", and the established mapping relationship is: {"Minority low-grade squamous intraepithelial lesion (LSIL)": "LSIL", "Low-grade squamous intraepithelial lesion (LSIL)": "LSIL", "No intraepithelial lesion or malignant cells found": "NILM", "No malignant cells found (NILM)": "NILM"}; for the "location" variable, the determined standard terms are "cervix", "vulva", "vagina", and "anus", and the established mapping relationship is: {"Posterior perineal commissure": "vulva", "vaginal wall": "vagina"}.

[0156] Step S326: Replace the original text variable values ​​with standard terminology.

[0157] The system iterates through the original text variable values, and based on the mapping table, replaces all corresponding original text variable values ​​with standard terms to complete the data normalization process.

[0158] In this embodiment, the aforementioned mapping dictionary is read, the original data columns are traversed, and all heterogeneous terms are replaced with standard terms. At this point, the original messy text data is transformed into standardized classification data. After the replacement, the variable values ​​for "TCT Result" are "LSIL", "LSIL", "HSIL", "SCC", "NILM", "NILM", and "ASC-H", while the variable values ​​for "Location" are "Cervix", "Vulva", "Vulva", "Vagina", "Anus", and "Vagina".

[0159] Step S330 involves classifying the data type of the rewritten text variables, including steps S332 to S334, where:

[0160] Step S332: Analyze whether there is a semantic-based hierarchical order among the standard terms.

[0161] This can be achieved by searching a pre-defined database of level keywords, such as "high / medium / low," "level," or "period," or by using a large language model to infer whether there is a progressive relationship between terms.

[0162] Step S334: Determine the order of levels:

[0163] If the hierarchical order exists, the data type of text variables is classified as ordered discrete.

[0164] When there is no ranking order, the data type of text variables is classified as unordered discrete.

[0165] In this embodiment, by combining the semantics of variable names and the characteristics of variable values, it is determined that "location" is an unordered discrete type. The data of "TCT results" contains sequential fields such as "low level" and "high level". Combining semantic information, it is determined that these six terms represent different stages of cervical epithelial lesions, and therefore it is determined to be an ordered discrete type.

[0166] This embodiment achieves standardized governance of unstructured text data by semantically clustering textual variables to identify heterogeneous words and establishing key-value mappings from heterogeneous words to standard terms based on context. This method can automatically identify and uniformly express words with the same meaning but different written forms, eliminating synonymous noise in the data. Furthermore, rewriting variable values ​​using standard terms not only standardizes the data format but also effectively reduces the cardinality of discrete variables, avoiding the curse of dimensionality in subsequent statistical models due to overly dispersed categories, and improving the analyzability of textual data.

[0167] In some embodiments, in step S140, the numerical agent classifies the numerical variables into continuous, unordered discrete, or ordered discrete types based on the numerical variable name and numerical variable value.

[0168] Numerical agents receive a sequence of pure numerical values, but do not treat it directly as a continuous sequence. Instead, they perform three-dimensional feature extraction and logical judgment, such as... Figure 4 As shown, the three-dimensional feature extraction includes semantic dimension analysis, value feature statistics, and hierarchical order detection.

[0169] Specifically, semantic dimension analysis includes natural language processing semantic analysis of numerical variable names, with semantics pointing to continuous change indicators, classification indicators, or hierarchical indicators; value feature statistics include scanning the distribution of numerical variable values, such as whether they contain decimals, or calculating the cardinality ratio, to determine whether the value features are continuous interval values ​​or discrete values; and hierarchy order detection is to detect whether variable values ​​have an inherent size order or logical hierarchy.

[0170] To determine if a data type is continuous, it needs to be determined whether the following three conditions are met simultaneously: the semantic state is an index pointing to continuous change; the value state is a continuous interval value; and the order state is none. If so, it is considered "continuous".

[0171] For example, if the variable is named "white blood cell count," its semantics point to a continuously changing indicator; the data shows a large number of decimals, such as 4.5, 4.51, and 10.2, meaning its value characteristic is a continuous range; and it has no hierarchical meaning, then the agent outputs a judgment result for this variable as continuous.

[0172] To determine whether an unordered discrete type is defined, it is necessary to check whether the following three conditions are met simultaneously: the semantic state points to a classification index; the value state is a discrete numerical value; and the order state is nonexistent. If so, it is identified as an "unordered discrete type".

[0173] For example, if the variable name is "gender," semantically referring to a classification indicator, and the data values ​​are discrete integers with no distinction between large and small, superior and inferior (no ranking order), the agent's output judgment result is unordered and discrete.

[0174] To determine whether an ordered discrete type is true, it is necessary to check whether the following three conditions are met simultaneously: the semantic state is a hierarchical index; the value state is a discrete numerical value; and the order state is present. If so, it is identified as an "ordered discrete type".

[0175] For example, if the variable name is "tumor staging" and the values ​​are only 1, 2, 3, and 4, the semantics refer to the grading index; and there is a clear relationship between the numerical values ​​in terms of disease severity (the order of the grades is yes). The agent outputs a judgment result that is an ordered discrete type.

[0176] If none of the above three combinations match, for example, if the semantics are categorical but the values ​​are continuous intervals, then mark it as awaiting manual review or force classification based on the value characteristics.

[0177] Based on the information in Table 1, the correspondence between the variable names and variable data types determined by the agent is as follows: "ID" is basic information; "Age" is continuous; "Marital status" is unordered discrete; "TCT result" is ordered discrete; "HPV result" is unordered discrete; "Location" is unordered discrete; and "Viral load" is continuous.

[0178] This embodiment utilizes a numerical agent to comprehensively analyze the semantics, value characteristics, and rank order of numerical variable names, achieving a fine-grained classification of the statistical attributes of numerical variables. It can accurately distinguish variable types that are physically stored as numbers but have completely different statistical meanings. For example, numbers representing classification codes are identified as unordered discrete, numbers representing rank scores are identified as ordered discrete, and numbers representing measurements are identified as continuous. This classification avoids meaningless mathematical operations such as averaging classification code values, ensuring that each type of variable can be matched with a governance logic that conforms to its statistical characteristics.

[0179] After pre-governance using intelligent agents, an intermediate multimodal governance dataset is obtained, as shown in Table 3.

[0180] Table 3

[0181]

[0182] Furthermore, based on the data types identified, corresponding statistical algorithms are executed to manage the data and generate managed data.

[0183] Outlier identification, missing value imputation, and standardization transformation are performed on continuous variables, which are numerical variables with continuous data type.

[0184] Specifically, for the "age" variable, it is initially identified as a numerical modality. The variable name semantically points to the time span, and the value characteristics are continuous integers without fixed levels. The statistical algorithm unit calls the outlier identification algorithm, using the interquartile range (IQR) algorithm to calculate the upper and lower bounds. For example, if "200 years old" is found to be an outlier, it is marked as "NA". The outlier identification algorithm can choose the three-standard-deviation method, the IQR method, or the absolute median difference method. The three-standard-deviation method is often used for variables whose original distribution follows a normal or approximately normal distribution. This algorithm considers that the area under the normal distribution curve within the interval μ±3σ can account for 99.7% of the total area under the normal distribution curve, where μ is the mean and σ is the standard deviation. If the data exceeds this range, it is reasonable to believe that its fluctuation is not only caused by random error, but may be due to non-random systematic error. Data outside the μ±3σ range can be identified as outliers. The interquartile range (IQR) method is often used for variables whose original distribution follows a skewed distribution. Data points below Q1 - 1.5 × IQR or above Q3 + 1.5 × IQR (where Q1 is the first quartile, IQR is the interquartile range itself, and Q3 is the third quartile) are considered outliers because their fluctuations are not solely due to random error but may involve non-random systematic errors. The absolute median difference (AMT) method is a more robust outlier identification algorithm; if data points meet the following conditions... , can be identified as outliers, among which For data points, For the median, Here, is the absolute median difference, and k is an empirical threshold. All three algorithms use common empirical thresholds and default parameters. Identified outliers are uniformly labeled "NA" for later imputation. It should be understood that no restrictions are placed on the outlier identification algorithm here.

[0185] For the variable "viral load", the initial values ​​included "not found", "1160 / vaginal", and "298000 cervical", which were identified as a mixed modality. The mixed data agent interpreted "not found" semantically as "no virus detected" and converted it to 0. For "1160 / vaginal", it used regular expressions to extract the value 1160 and discarded the irrelevant text " / vaginal". For "298000 cervical", it extracted the value 298000. However, the agent subsequently recognized a large range in the variable's value distribution, from 0 to 298000, therefore, 298000 was identified as an outlier by the statistical algorithm and marked as missing.

[0186] For imputation of missing values ​​in continuous numerical variables, this application employs a strategy of parallel imputation using multiple algorithms, each generating an imputed dataset. The algorithms include mean-mode imputation, K-nearest neighbor imputation, linear regression imputation, and multiple imputation, specifically:

[0187] In the mean-mode imputation method, the kurtosis value of the continuous variable is first calculated to determine whether its distribution type is normal or skewed; if it is normal, the arithmetic mean is used for imputation, and if it is skewed, the median is used for imputation.

[0188] In the K-nearest neighbor imputation method, similar samples without missing values ​​are selected based on the Gaussian distance, the mean of the continuous variable to be imputed in the sample is calculated, and the continuous variable is imputed.

[0189] In linear regression imputation, a generalized linear regression model is constructed in samples where the continuous variable to be imputed has no missing values, and imputed values ​​for the continuous variable are generated based on this model.

[0190] In the multiple imputation method, a predictive mean matching algorithm is used to iteratively generate multiple imputed datasets for continuous variables to ensure robustness, and multiple imputed datasets are retained simultaneously.

[0191] All imputation algorithms have preset default parameters, requiring no manual intervention and adapting to fully automated processing. After running the above four imputation algorithms, a total of 4+m imputed datasets are generated, where m refers to the preset number of datasets to be imputed in the multiple imputation method. Except for the mean-mode imputation method, in actual processing, there may be extreme cases where imputation fails due to a large proportion of missing data. In this case, the mean-mode imputation method is selected as the second choice to generate the imputed value for the variable.

[0192] After imputing missing values, a standardization transformation is performed. The entire data column is standardized using methods such as natural logarithm transformation, normalization, or minimum-maximum normalization to ensure it conforms to a standard normal distribution and avoids the influence of extreme values. In some embodiments, the original values ​​are also retained simultaneously to adapt to different needs. Specifically, the transformation formulas used include, but are not limited to:

[0193]

[0194] Where x is the original value, x ’ The values ​​are standardized, μ is the mean of the original values, σ is the standard deviation of the original values, and x is the value after standard transformation. i Let x be the i-th original value, min(x) be the minimum value in the original value data, and max(x) be the maximum value in the original value data.

[0195] Missing value imputation and small sample class merging are performed on unordered discrete variables, which include text variables and numerical variables of unordered discrete data type.

[0196] For missing value imputation of unordered discrete variables (including textual and numerical unordered discrete variables), similar to the aforementioned continuous variables, a strategy of parallel imputation using multiple algorithms is employed. Each algorithm generates an imputed dataset. The algorithms include mean-mode imputation, K-nearest neighbor imputation, linear regression imputation, and multiple imputation. Here, only the adaptation adjustment is described:

[0197] In the mean-mode imputation method, the most frequent class is selected for mode imputation; in the K-nearest neighbor imputation method, neighboring samples are selected based on class similarity, and the most frequent class is used for imputation; in the linear regression imputation method, unordered discrete variables need to be one-hot encoded for classes before modeling, and this method is only for supplementary adaptation; in the multiple imputation method, unordered discrete variables iteratively predict class probabilities and select the optimal class for imputation. All imputation processes are completed automatically through preset algorithm logic without manual intervention, ensuring governance efficiency and consistency.

[0198] For the unordered discrete variables "location", "marital status" and "HPV result", a small sample category merging operation is performed; categories whose number of samples or whose sample proportion is less than a preset threshold are merged and marked as "other" to distinguish them from the remaining categories and reduce model noise.

[0199] Missing value imputation and small sample adjacent category merging operations are performed on ordered discrete variables. Ordered discrete variables include text variables and numerical variables with ordered discrete data types.

[0200] For missing value imputation of ordered discrete variables, similar to the strategy for continuous variables, a parallel imputation strategy using multiple algorithms is employed. Each algorithm generates an imputed dataset. The algorithms include mean-mode imputation, K-nearest neighbor imputation, linear regression imputation, and multiple imputation. Here, only the adaptation adjustment is described:

[0201] In the mean-mode imputation method, the mode of the hierarchy or the hierarchy that matches the business rules is selected for imputation, preserving the hierarchy order logic; in the K-nearest neighbor imputation method, similarity is calculated by combining hierarchy weights, and the core hierarchy of neighboring samples is selected for imputation first; in the linear regression imputation method, the hierarchy is first ordered and encoded, for example, 1=low, 2=medium, 3=high, and then a regression model is constructed to ensure that the hierarchy logic is not destroyed; in the multiple imputation method, the hierarchy order constraint is strictly followed during the iteration of ordered discrete variables to avoid imputation across levels.

[0202] For example, when imputing missing values ​​for ordered discrete variables, the system first checks whether there are pre-defined hierarchical imputation rules for the variable in the variable dictionary. If so, it directly replaces the missing values ​​with the corresponding hierarchical categories according to the rules, which conforms to the inherent hierarchical order of the variable and does not exceed the hierarchical range.

[0203] If the variable dictionary does not have relevant rules, the hierarchical mode imputation method is used: count the hierarchical frequency of non-missing values ​​of the variable, select the hierarchical mode with the highest frequency, and replace all missing values ​​with the hierarchical mode; if there are multiple hierarchical modes, the middle category in the hierarchical sequence or the first occurrence of the hierarchical category is selected by default.

[0204] If the proportion of missing values ​​is greater than the preset threshold and the hierarchical distribution is uniform, the adjacent hierarchical interpolation imputation method is used: select the K most valid samples that are closest to the missing samples, and take the hierarchical level with the highest frequency or the middle level as the imputation value.

[0205] Subsequently, a small sample adjacent category merging operation is performed on the ordered discrete variable "TCT result". Categories with fewer than a preset number threshold or a sample ratio less than a preset ratio threshold are merged into adjacent level categories according to the semantic recognition level order.

[0206] Specifically, firstly, the large model identifies the semantics of the variables and sorts the ordered variables from smallest to largest level; secondly, it counts missing values ​​for all levels, identifies the first small sample class from right to left, and merges it with the adjacent left-hand class; the previous step is repeated until the leftmost class is a non-small sample class; finally, the leftmost non-small sample class is merged with all the small sample classes to its left.

[0207] For example, for ordered discrete variables such as "satisfaction level: low, medium, high", an adjacency category merging operation is performed. If very few samples of "low" satisfaction are found, they will not be merged into "high". Instead, based on the semantic level of adjacency, they will be merged into the nearest "medium" category, or "very low" will be merged into "low" to protect the ordinal logic of the data from being violated.

[0208] Evaluate the data governance effectiveness under all combinations of governance schemes, including whether the post-governance dataset contains any unremoved potential personal information data, unidentified missing representations or unfilled missing values, arbitrary duplicate rows or columns, constant columns, and arbitrary small sample classes for discrete variables. If so, provide the specific variable names and small sample classes. Compare the dataset quality before and after governance, using missing data ratio heatmaps, continuous variable box plots, or discrete variable class percentage plots for measurement and evaluation.

[0209] This embodiment performs differentiated governance operations based on the finely categorized data types, significantly optimizing the quality of the data distribution after governance. For continuous variables, outlier identification and standardization are performed, eliminating dimensional differences and extreme value interference. For unordered discrete variables, small sample categories are merged, enhancing the model's generalization ability. In particular, for ordered discrete variables, adjacent category merging is performed, addressing the sample sparsity problem while strictly preserving the inherent ranking information of the variables, preventing the destruction of potential data trends due to arbitrary category merging, thereby maximizing the preservation of the data's business meaning and predictive value.

[0210] By constructing a dual verification method that combines semantic features of variable names with statistical features of numerical distribution, the essential attributes of mixed data can be quantitatively determined through technical means. For variables with a numerical modality, a character stripping algorithm is automatically executed to extract valid values; for variables with a text modality, the original values ​​are automatically locked. This overcomes the information distortion problem caused by traditional single-rule processing of mixed data, significantly improving the accuracy of data cleaning.

[0211] Secondly, this application enhances the data governance system's ability to handle specific data. By configuring the variable dictionary mapping interface, pre-defined static imputation rules can be integrated into the dynamic governance process. This allows for priority matching of defined rule parameters when handling missing values, effectively avoiding logical conflicts of conventional statistical algorithms in specific business scenarios without increasing computational overhead, thus ensuring the compliance of data imputation.

[0212] Finally, this application improves the robustness and error correction capabilities of the data governance system. By setting up a result verification interaction interface, it can receive feedback signals from external inputs and correct and update the internally generated data type labels based on these signals. By adding a correction step, it ensures that the data labels output to the subsequent statistical modules match the actual data content, thereby guaranteeing the accuracy and execution efficiency of subsequent statistical algorithm calls.

[0213] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0214] Based on the same inventive concept, this application also provides an apparatus for implementing the multimodal data governance method based on multi-agent cooperation described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method. Therefore, the specific limitations of one or more embodiments of the multimodal data governance apparatus 500 based on multi-agent cooperation provided below can be found in the limitations of the multimodal data governance method based on multi-agent cooperation described above, and will not be repeated here.

[0215] In an exemplary embodiment, a multimodal data governance device 500 based on multi-agent collaboration is provided, comprising: a data exploration module 510, a hybrid processing module 520, a text processing module 530, a numerical processing module 540, and a data cleaning module 550, wherein:

[0216] The data exploration module 510 is used to acquire the data to be governed, scan the initial variable values ​​of each variable in the data to be governed, and identify the variable modality of the variable as text modality, numerical modality or mixed modality based on the distribution of the initial variable values.

[0217] The hybrid processing module 520 is used to redefine and transform hybrid variables belonging to hybrid modalities using a hybrid data agent, converting them into text modal variables or numerical modal variables;

[0218] The text processing module 530 is used to obtain text-type variables whose variable modality is text mode. The text-type variables include variables whose variable modality is identified as text mode and the converted text-type variables. The text agent is used to perform standardized evaluation and corresponding processing on the text-type variables, and to classify the data type of the text-type variables into non-standard text type, unordered discrete type or ordered discrete type.

[0219] The numerical processing module 540 is used to acquire numerical variables whose variable modes are numerical modes. Numerical variables include variables whose variable modes are identified as numerical modes and the converted numerical mode variables. The numerical agent is used to classify the data type of numerical variables into continuous, unordered discrete, or ordered discrete types based on the numerical variable name and numerical variable value.

[0220] The data cleaning module 550 is used to perform corresponding statistical algorithm processing operations based on the data types to generate processed data.

[0221] like Figure 5 As shown, the output of the data exploration module 510 is connected to the input of the hybrid processing module 520, the input of the text processing module 530, and the input of the numerical processing module 540, respectively.

[0222] The output of the hybrid processing module 520 is connected to the input of the text processing module 530 and the input of the numerical processing module 540, respectively. The hybrid processing module 520 is used to send the text modal variables obtained by the hybrid modal conversion to the text processing module 530 and send the numerical modal variables obtained by the hybrid modal conversion to the numerical processing module 540.

[0223] The output of the text processing module 530 is connected to the input of the data cleaning module 550. The text processing module 530 is used to send text variables that have been rewritten according to standard terms and divided into unordered discrete or ordered discrete types to the data cleaning module 550.

[0224] The output of the numerical processing module 540 is connected to the input of the data cleaning module 550. The numerical processing module 540 is used to send numerical variables classified as continuous, unordered discrete, or ordered discrete to the data cleaning module 550.

[0225] It should be noted that the specific processes, technical details, and extended functions of each module / unit performing operations in the above-described embodiments of the multimodal data governance device 500 based on multi-agent collaboration have been described in detail and correspondingly in the embodiments of the multimodal data governance method based on multi-agent collaboration. This device can implement all the steps and functions provided by any of the aforementioned method embodiments and achieve the same technical effect. To avoid repetition and to limit the length of the specification, further details are omitted here.

[0226] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 6As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When executed by the processor, the computer program implements a multi-agent collaborative multimodal data governance method. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0227] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0228] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0229] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0230] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0231] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0232] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0233] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0234] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A multimodal data governance method based on multi-agent collaboration, characterized in that, The method includes: Acquire the data to be governed, scan the initial variable values ​​of each variable in the data to be governed, and identify the variable modality of the variable as text modality, numerical modality or mixed modality based on the value distribution of the initial variable values; The hybrid data agent is used to redefine and transform the hybrid variables belonging to the hybrid modality, converting them into text modal variables or numerical modal variables; Obtain text-type variables whose variable modality is the text modality, wherein the text-type variables include variables whose variable modality is identified as text modality and the text modality variables obtained by conversion; The text-type variables are standardized and processed using a text agent, including: analyzing the distribution pattern of the text variable values ​​of all the text-type variables; if the distribution pattern is not statistically regular, the standardization evaluation result is determined to be that standardization is not required, and the data type of the text-type variables is classified as non-standard text type. If the distribution pattern exhibits statistical regularity, the standardized evaluation result is determined to require standardization. Based on the textual variables, a standard terminology mapping is generated, and the textual variable values ​​are rewritten. The rewritten textual variables are then classified into unordered discrete or ordered discrete data types, where: The process of generating standard terminology mappings and rewriting text variable values ​​includes performing semantic clustering on the text variable values, establishing a mapping relationship between heterogeneous words and standard terms, and replacing the original text variable values ​​with the standard terms. The classification of data types for the rewritten text variables includes: Analyze whether there is a semantic-based hierarchical order among the standard terms; If the specified hierarchical order exists, the data type of the text variable is classified as ordered discrete. When the hierarchical order is zero, the data type of the text variable is classified as unordered discrete. Obtain numerical variables whose variable mode is the numerical mode, wherein the numerical variables include variables whose variable mode is identified as numerical modes and the numerical mode variables obtained by conversion; Using a numerical agent, based on the numerical variable name and numerical variable value of the numerical variable, the semantics of the numerical variable name is analyzed, and the value characteristics and rank order of the numerical variable value are statistically analyzed. In the semantically oriented, continuously changing index, where the value feature is a continuous interval value and the level order is absent, the numerical variable is classified as continuous. In the semantically oriented, classification index, where the value feature is a discrete value and the level order is absent, the numerical variable is classified as unordered discrete. In the semantically oriented, hierarchical index, where the value feature is a discrete value and the level order is present, the numerical variable is classified as ordered discrete. Based on the data type classification, the corresponding statistical algorithm is executed to perform the data management operation and generate the managed data.

2. The method according to claim 1, characterized in that, The step of scanning the initial variable values ​​of each variable in the data to be treated, and identifying the variable modality of the variable as text modality, numerical modality, or mixed modality based on the distribution of the initial variable values, includes: Traverse the initial variable values ​​of each variable in the data to be managed to identify character distribution characteristics; Based on the character distribution characteristics, determine the variable mode of the variable: If the initial variable values ​​are all composed of numeric characters, the variable mode of the corresponding variable is determined to be a numeric mode. If the initial variable values ​​are all composed of non-numeric characters, the variable modality of the corresponding variable is determined to be text modality; If the initial variable value contains both numeric characters and non-numeric characters, the variable mode of the corresponding variable is determined to be a mixed mode.

3. The method according to claim 1, characterized in that, The process of redefining and transforming hybrid variables belonging to the hybrid modality using a hybrid data agent includes: Read the mixed variable name and mixed variable value of the mixed type variable; Identify the semantic meaning of the mixed variable names and the distribution characteristics of the numerical content in the mixed variable values; Based on the semantic orientation and the distribution characteristics, the hybrid variable is redefined as a text modal variable or a numerical modal variable; When redefined as a text modal variable, the original content of the mixed variable value is retained as the text variable value of the transformed text modal variable; When redefined as a numerical modal variable, the numerical content of the mixed variable value is extracted as the numerical variable value of the transformed numerical modal variable.

4. The method according to claim 1, characterized in that, The step of performing corresponding statistical algorithm governance operations based on the data types to generate governance data includes: Outlier identification, missing value imputation, and standardization transformation are performed on the continuous variables, where the continuous variables are numerical variables whose data type is continuous. Missing value imputation and small sample class merging operations are performed on the unordered discrete variables. The unordered discrete variables include text variables and numerical variables whose data type is unordered discrete. Missing value imputation and small sample adjacent category merging operations are performed on the ordered discrete variables. The ordered discrete variables include text variables of the data type ordered discrete and numerical variables of the data type ordered discrete.

5. A multimodal data governance device based on multi-agent collaboration, characterized in that, The device includes: The data exploration module is used to acquire the data to be managed, scan the initial variable values ​​of each variable in the data to be managed, and identify the variable modality of the variable as text modality, numerical modality or mixed modality according to the value distribution of the initial variable values. The hybrid processing module is used to redefine and transform hybrid variables belonging to the hybrid modality using a hybrid data agent, converting them into text modality variables or numerical modality variables; A text processing module is used to acquire text-type variables whose variable modality is the text modality, the text-type variables including variables whose variable modality is identified as text modality and the converted text modality variables; and to perform standardization evaluation and corresponding processing on the text-type variables using a text agent, including: analyzing the distribution pattern of the text variable values ​​of all the text-type variables; and if the distribution pattern is not statistically regular, determining that the standardization evaluation result does not require standardization, and classifying the data type of the text-type variables as non-standard text type. If the distribution pattern exhibits statistical regularity, the standardized evaluation result is determined to require standardization. Based on the textual variables, a standard terminology mapping is generated, and the textual variable values ​​are rewritten. The rewritten textual variables are then classified into unordered discrete or ordered discrete data types, where: The process of generating standard terminology mappings and rewriting text variable values ​​includes performing semantic clustering on the text variable values, establishing a mapping relationship between heterogeneous words and standard terms, and replacing the original text variable values ​​with the standard terms. The classification of data types for the rewritten text variables includes: Analyze whether there is a semantic-based hierarchical order among the standard terms; If the specified hierarchical order exists, the data type of the text variable is classified as ordered discrete. When the hierarchical order is zero, the data type of the text variable is classified as unordered discrete. The numerical processing module is used to obtain numerical variables whose variable mode is the numerical mode, the numerical variables including variables whose variable mode is identified as numerical mode and the numerical mode variables obtained by conversion; and to use a numerical agent to analyze the semantics of the numerical variable names and numerical variable values ​​based on the numerical variable names and numerical variable values ​​of the numerical variables, and to statistically analyze the value characteristics and rank order of the numerical variable values. In the semantically oriented, continuously changing index, where the value feature is a continuous interval value and the level order is absent, the numerical variable is classified as continuous. In the semantically oriented, classification index, where the value feature is a discrete value and the level order is absent, the numerical variable is classified as unordered discrete. In the semantically oriented, hierarchical index, where the value feature is a discrete value and the level order is present, the numerical variable is classified as ordered discrete. The data cleaning module is used to perform corresponding statistical algorithm processing operations based on the data types to generate processed data.

6. The apparatus according to claim 5, characterized in that, The step of scanning the initial variable values ​​of each variable in the data to be managed, and identifying the variable modality of the variable as text modality, numerical modality, or mixed modality based on the distribution of the initial variable values, includes: Traverse the initial variable values ​​of each variable in the data to be managed to identify character distribution characteristics; Based on the character distribution characteristics, determine the variable mode of the variable: If the initial variable values ​​are all composed of numeric characters, the variable mode of the corresponding variable is determined to be a numeric mode. If the initial variable values ​​are all composed of non-numeric characters, the variable modality of the corresponding variable is determined to be text modality; If the initial variable value contains both numeric characters and non-numeric characters, the variable mode of the corresponding variable is determined to be a mixed mode.

7. The apparatus according to claim 5, characterized in that, The process of redefining and transforming hybrid variables belonging to the hybrid modality using a hybrid data agent includes: Read the mixed variable name and mixed variable value of the mixed type variable; Identify the semantic meaning of the mixed variable names and the distribution characteristics of the numerical content in the mixed variable values; Based on the semantic orientation and the distribution characteristics, the hybrid variable is redefined as a text modal variable or a numerical modal variable; When redefined as a text modal variable, the original content of the mixed variable value is retained as the text variable value of the transformed text modal variable; When redefined as a numerical modal variable, the numerical content of the mixed variable value is extracted as the numerical variable value of the transformed numerical modal variable.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Variant language violation text recognition method, system and equipment and medium

    CN121016206A

  • Methods, systems, and computer program products for generating 3D human pose and movement estimation from monocular image information

    US20260065565A1