Data processing method and device, electronic equipment and storage medium
By using a cube model to model the subjects, tasks, and processing stages as a three-dimensional structure for queue data modeling, the problems of storage space waste and data retrieval in existing technologies are solved, and efficient data management and statistics are achieved.
Patent Information
- Application Number
- CN202210272804.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-18
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2042-03-18
AI Technical Summary
Existing technologies use dimensions as the basic element for cohort data modeling, which leads to a waste of storage space in longitudinal cohort studies and increases the difficulty of retrieving and computing multimodal data, making it difficult to meet the macro-management needs of subjects and data.
Using a cube model as the basic element, the cohort data is modeled. The cohort data is generated using the subjects, tasks, and processing stages as a three-dimensional structure, and information retrieval and statistics are performed through a retrieval platform.
It effectively avoids wasting storage space, improves data processing efficiency, meets the macro-statistical needs of multimodal data in longitudinal cohort studies, and improves data retrieval and query efficiency.
Smart Images

Figure CN114936196B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a data processing method, apparatus, electronic device, and storage medium. Background Technology
[0002] Longitudinal cohort studies are a common research method in epidemiology, and an effective way to explore pathogenic risk factors and evaluate interventions. Cohort data is one of the core outcomes of longitudinal cohort studies, and its quality directly determines the success or failure of the study.
[0003] Typically, cohort data share the following characteristics: (1) the data modalities are diverse, often involving multimodal and multiscale data; (2) the subjects are dynamically variable, and during long-term follow-up, some subjects may drop out, and new subjects may be added; (3) the testing tasks are dynamically variable, and the specific tasks may differ when collecting data of the same modality from the same group at different times.
[0004] The aforementioned characteristics bring complex situations to the management of cohort data, especially in the stage of cohort data modeling applications. Existing technologies for multi-dimensional cohort data use dimensions as basic elements for modeling, which leads to a waste of storage space and increases the difficulty of retrieval and calculation of multimodal data of subjects, making it difficult to meet the needs of longitudinal cohort research to focus on subjects and macro-management of data. Summary of the Invention
[0005] This invention provides a data processing method, apparatus, electronic device, and storage medium to address the shortcomings of existing technologies that use dimensions as basic elements for queue data modeling, which affects the effectiveness of longitudinal queue research.
[0006] This invention provides a data processing method, comprising:
[0007] Based on multimodal data, cohort data is generated. The multimodal data is obtained by collecting data on tasks in at least two task rounds at at least two stations during the cohort study. The multimodal data includes data from at least two modalities, and each station includes at least two subjects.
[0008] Using a cube model as the basic element, data modeling is performed on the queue data to obtain modeling data. The cube model is a three-dimensional structure with subjects, tasks, and processing stages as the three dimensions.
[0009] According to a data processing method provided by the present invention, generating queue data based on multimodal data includes:
[0010] Based on the tasks corresponding to the data of the at least two modalities and the task rounds to which the tasks belong, and / or the subjects corresponding to the data of the at least two modalities and the stations to which the subjects belong, the at least two modal data are integrated to obtain the queue data.
[0011] According to a data processing method provided by the present invention, the step of using a cube model as the basic element to perform data modeling on the queue data to obtain modeled data includes:
[0012] Using the cube model as the basic element, and based on the tasks and task rounds corresponding to the modalities of each basic data in the queue data, and / or the subjects and processing stages corresponding to each basic data, modeling is performed to obtain the modeling data;
[0013] The data structure of the modeling data is a three-dimensional structure with three dimensions: site structure, task round structure, and processing stage. The site structure is a hierarchical structure including sites and subjects under each site, and the task round structure is a hierarchical structure including task rounds and tasks under each task round.
[0014] According to a data processing method provided by the present invention, the step of using a cube model as the basic element to perform data modeling on the queue data to obtain modeled data further includes:
[0015] Determine the target parameters for the dimension to be statistically analyzed;
[0016] Based on the position of the target parameter in the dimension to be counted in the modeling data, planar data that conforms to the target parameter is determined from the modeling data;
[0017] Based on the aforementioned planar data, statistical analysis is performed.
[0018] According to a data processing method provided by the present invention, when the dimension to be statistically analyzed is a task, the step of performing data statistics based on the planar data includes:
[0019] Based on the projection length of the planar data in the direction corresponding to the subject, the amount of data for the task corresponding to the target parameter is determined.
[0020] And / or,
[0021] The data in the planar data whose processing stage is the data acquisition stage is identified as acquired data, and the data in the planar data whose processing stage is the quality control stage is identified as quality control data.
[0022] Based on the projection length of the collected data in the direction corresponding to the subject, and the projection length of the quality control data in the direction corresponding to the subject, the quality control progress of the task corresponding to the target parameter is determined.
[0023] According to a data processing method provided by the present invention, when the dimension to be statistically analyzed is the subject, the step of performing data statistics based on the planar data includes:
[0024] Determine the current task round;
[0025] Based on the projection length of the planar data in the direction corresponding to the task, the task round of the latest task corresponding to the target parameter is determined;
[0026] Remove the target parameters of the task rounds in which the latest task is located before the current round, and perform subject statistics based on the remaining target parameters.
[0027] According to a data processing method provided by the present invention, the step of using a cube model as the basic element to perform data modeling on the queue data to obtain modeled data further includes:
[0028] Receive query information;
[0029] Based on the retrieval platform, the query information is retrieved to obtain the retrieval results;
[0030] The retrieval platform is built based on the modeling data, as well as information toolsets and / or project information, which are stored in a structure that includes at least two dimensions: tasks and processing phases.
[0031] According to a data processing method provided by the present invention, the at least two modalities include at least two of cognitive psychological behavior, magnetic resonance imaging, electroencephalography, biological samples, and the environment;
[0032] The processing phase includes at least two of the following: data acquisition phase, quality control phase, analysis phase, and sharing phase.
[0033] According to a data processing method provided by the present invention, the quality control stage includes a site quality control stage and a central quality control stage. The site quality control stage is used to perform quality control on data collected from the same site, and the central quality control stage is used to perform quality control on data collected from at least two sites.
[0034] The present invention also provides a data processing apparatus, comprising:
[0035] The data determination unit is used to generate cohort data based on multimodal data. The multimodal data is obtained by collecting data on tasks in at least two task rounds at at least two stations during the cohort study. The multimodal data includes data from at least two modalities, and each station includes at least two subjects.
[0036] The data modeling unit is used to perform data modeling on the queue data based on the cube model as the basic element, so as to obtain modeling data. The cube model is a three-dimensional structure with subjects, tasks and processing stages as the three dimensions.
[0037] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement any of the data processing methods described above.
[0038] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the data processing method as described above.
[0039] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements any of the data processing methods described above.
[0040] The data processing method, apparatus, electronic device, and storage medium provided by this invention are based on a cube model with subjects, tasks, and processing stages as three dimensions. This modeling method performs data modeling on data obtained from cohort studies and processes the modeled data. This effectively avoids the problem of wasted storage space, solves the business needs of macro-statistics of multimodal data in longitudinal cohort studies, and helps to improve data processing efficiency. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0042] Figure 1 This is a flowchart illustrating the data processing method provided by the present invention;
[0043] Figure 2 This is a schematic diagram of the modeling data provided by the present invention;
[0044] Figure 3 This is a schematic diagram of the site structure provided by the present invention;
[0045] Figure 4 This is a schematic diagram of the task round structure provided by the present invention;
[0046] Figure 5 This is a flowchart illustrating the modeling data processing method provided by the present invention;
[0047] Figure 6 This is a schematic diagram of the structure of the multimodal data provided by the present invention;
[0048] Figure 7 This is a schematic diagram of a vertical queue scenario provided by the present invention;
[0049] Figure 8 This is a schematic diagram of the data processing device provided by the present invention;
[0050] Figure 9 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0052] Longitudinal cohort studies are a common research method in epidemiology. They are an effective way to explore pathogenic risk factors and evaluate interventions, and are gradually becoming an important research tool for discovering patterns and exploring mechanisms in fields such as clinical precision medicine, chronic disease management, and evidence-based education.
[0053] When modeling data for cohort data, related technologies consider the multi-dimensional nature of cohort data. The modeling method mainly uses the data metric of a single modality as the modeling object to complete the statistical analysis of specific data modalities. If the dimensional metrics obtained from two modalities are different, it is difficult to perform joint modeling. If the dimensional metrics between two or more modalities are linearly spliced, it will cause strong sparsity in the constructed cube model. This not only wastes the platform's data storage space, but also hinders the efficient response to the multi-modal data retrieval and calculation needs of the subjects. As a result, it is difficult to meet the needs of longitudinal cohort studies to focus on subjects and macro-manage data.
[0054] To address this problem, the present invention provides a data processing method. Figure 1 This is a flowchart illustrating the data processing method provided by the present invention, as shown below. Figure 1 As shown, the method includes:
[0055] Step 110: Generate queue data based on multimodal data. The multimodal data is obtained by collecting data on tasks in at least two task rounds at at least two sites during the queue study. The multimodal data includes data from at least two modalities, and each site includes at least two subjects.
[0056] Step 120: Using the cube model as the basic element, perform data modeling on the queue data to obtain modeling data. The cube model is a three-dimensional structure with subjects, tasks, and processing stages as the three dimensions.
[0057] Specifically, multimodal data refers to data generated in longitudinal cohort studies. The specific data included in multimodal data can be determined based on the content of the longitudinal cohort study. For example, when studying the brain development of a population, multimodal data can include at least two of the following: cognitive psychological behavior data, magnetic resonance imaging data, electroencephalogram (EEG) data, and biological sample data.
[0058] Considering that longitudinal cohort studies are characterized by multiple sites (geographic dimension), multimodal (data dimension), long duration (tracking for many years), full process (data collection, quality control, analysis, and sharing), and large scale (data volume dimension), if data modeling is performed from only a single dimension based on relevant technologies, the resulting data scale will inevitably be extremely large, resulting in wasted storage space and affecting data retrieval and computation efficiency. To address this, this invention formally defines cohort data from three dimensions: subjects, tasks, and processing stages. This allows for the setting of basic data modeling elements in a cubic form within a three-dimensional space, namely, the smallest granularity data cube model. Based on this, cohort data can be integrated from multimodal data, and then modeling can be performed based on the cohort data to obtain modeling data.
[0059] Longitudinal cohort studies are characterized by multiple sites, meaning that the samples in the cohort often come from a geographically dispersed area. Because the multimodal data collected in cohort studies are limited by experimental instruments and equipment, at least two research institutions / hospitals with the necessary experimental capabilities are typically selected as data collection points; these data collection points are called sites. Given the consideration of geographical representativeness of the samples in cohort studies, a multi-site design is usually adopted, meaning a set of multiple sites is given for the cohort study.
[0060] In actual cohort studies, a site can include two types of personnel: subjects and experimenters. Subjects refer to the participants in the cohort study, who are the contributors and data collectors. Experimenters refer to the implementers of the cohort study, who are the operators of a series of processes such as data collection and processing. Considering that subjects reflect the data collectors in the cohort, and that subjects can appear as personnel within a site, for multimodal data collected from multiple sites, cohort data can be generated based on the sites and subjects within those sites. Subjects can be used as one dimension of the basic elements, reflecting both the subjects and the sites from which the data originates in the data model.
[0061] Longitudinal cohort studies are characterized by multimodality, meaning that data collection from participants may involve at least two modalities, such as cognitive, psychological, and behavioral data, magnetic resonance imaging (MRI) data, electroencephalogram (EEG) data, and biological sample data. The resulting multimodal data can contain data from each modality.
[0062] Longitudinal cohort studies are characterized by their long duration, meaning that data collection may involve multiple waves. A longitudinal cohort study is a type of research that starts at a specific point in time and collects data from participants continuously for several years or at intervals of several years. Each data collection session is called a wave, and the duration of a wave can be 12-18 months or other lengths. Each wave may also involve data collection from various modalities.
[0063] Whether data collection is conducted across multiple rounds or within a single round, the collection of data for each modality can be considered a data collection task. Therefore, when generating queue data based on multimodal data, the data for each modality acquired in each task round can be mapped to the corresponding task in that round. This means leveraging the mapping between modalities and tasks to associate the data of each modality within the queue with the tasks. Based on this, in data modeling, tasks can be considered as a dimension within the basic elements. Representing the tasks associated with the data reflects both the data's modality and the round in which it is collected.
[0064] The characteristic of a longitudinal cohort study's entire process (phases) is that data processing in a longitudinal cohort study can include data acquisition, data quality control, data analysis, and data sharing, which together constitute the entire data processing flow. Therefore, when generating cohort data based on multimodal data, the processing stage of each data point within the cohort can be determined. Based on this, in data modeling, the processing stage can be used as a dimension among the basic elements to reflect the processing stage of the data.
[0065] In summary, the subjects, tasks, and treatment phases can be used as the three dimensions of a cube model to model cohort data, thereby obtaining modeling data that can also be represented from the three dimensions of subjects, tasks, and treatment phases.
[0066] Furthermore, the three dimensions of the modeling data here can be reflected as three coordinate axes in a three-dimensional coordinate system. For example, the X-axis represents the subject, the Y-axis represents the task, and the Z-axis represents the data processing phase. Data for a specific modality can be stored in a defined 1×1×1 granularity unit cube model in three-dimensional space. For example, Figure 2This is a schematic diagram of the modeling data provided by the present invention. Figure 2 In this process, the usage phase can be divided into the data acquisition phase, the local quality control phase, the central quality control phase, the data analysis phase, and the data sharing phase. Tasks can be divided into Behavior 1, ..., Magnetic Resonance 1, Behavior 2, etc., where Behavior represents the acquisition task corresponding to cognitive psychological behavioral data, Magnetic Resonance represents the acquisition task corresponding to magnetic resonance data, and 1, 2, etc. represent the round in which the task is located. For example, Behavior 1 is the acquisition task corresponding to cognitive psychological behavioral data in round 1, Behavior 2 is the acquisition task corresponding to cognitive psychological behavioral data in round 2, and Behavior 1 occurs before Behavior 2. In addition, the order of different tasks in the same round, such as Behavior 1 and Magnetic Resonance 1, is not used to limit the execution order. Figure 2 The cubes filled with different filling methods can represent the magnetic resonance data of subject 2 at each stage of use in round 1.
[0067] Modeling data obtained based on a cube model can reflect relevant information from three dimensions. Compared to modeling based on only a single dimension, it can greatly compress the data size. As a result, the large scale of the queue data itself can be reasonably reduced, avoiding waste of storage space and improving the efficiency of subsequent data processing.
[0068] The method provided in this invention is based on a cube model with subjects, tasks, and processing stages as three dimensions to model the data obtained from cohort studies. This effectively avoids the problem of wasted storage space, solves the business needs of macro-statistics of multimodal data in longitudinal cohort studies, and helps to improve data processing efficiency.
[0069] Based on the above embodiments, step 110 includes:
[0070] Based on the tasks corresponding to the data of the at least two modalities and the task rounds to which the tasks belong, and / or the subjects corresponding to the data of the at least two modalities and the stations to which the subjects belong, the at least two modal data are integrated to obtain the queue data.
[0071] Specifically, the multimodal data generated based on longitudinal cohort studies can cover data from at least two modalities. The processing of multimodal data involves information such as the data collection sites, the task rounds of data collection, the specific tasks under the task rounds, and the processing stages of data processing. Therefore, when integrating the data of each modality in the multimodal data to obtain cohort data that can be used for modeling, the data of each modality can be integrated from the tasks corresponding to each modality and the task rounds to which the tasks belong, and / or the subjects corresponding to each modality and the sites to which the subjects belong, in order to generate cohort data.
[0072] The queue data is generated based on the timing of each participant's task execution at each station and in each task round. For example, the generated queue data can be represented in the following form:
[0073]
[0074] in This indicates the collection of data for a specific task within a specific modality, i.e. It can be data from the i-th task in a cognitive psychology and behavioral modality. Or the data of the i-th task in the magnetic resonance mode. Or the data from the i-th task in the EEG modality Or data from the i-th task in the biological sample modality. Or the data of the i-th task in the environment variable modality. Wait, w l express Belongs to the first task round, sub ki express For the i-th subject at the k-th site, exp kj express For the j-th subject under the k-th site, s k express For the k-th station, t represents the time when the task is started.
[0075] It should be noted that the generation of queue data occurs during the data acquisition phase, while the remaining processing phases, such as the quality control phase, analysis phase, and sharing phase, are all implemented after the queue data is generated.
[0076] Based on any of the above embodiments, step 120 includes:
[0077] Using the cube model as the basic element, and based on the tasks and task rounds corresponding to the modalities of each basic data in the queue data, as well as the subjects and processing stages corresponding to each basic data, modeling is performed to obtain the modeling data.
[0078] The data structure of the modeling data is a three-dimensional structure with site structure, task round structure and processing stage as the three dimensions.
[0079] The site structure is a hierarchical structure including sites and subjects under each site, and the task round structure is a hierarchical structure including task rounds and tasks under each task round.
[0080] Specifically, when performing data modeling based on queue data, corresponding to the three dimensions of the cube model, the modeling of the queue data needs to be based on the task corresponding to the modality of the basic data, the subject corresponding to the basic data, and the processing stage corresponding to the basic data. Here, the basic data refers to the data obtained by dividing the queue data into basic elements using the cube model as the basic element. Each basic data corresponds to the data of a subject collected from a processing stage in a task.
[0081] The correspondence between modalities and tasks is pre-defined. Considering that the data of each modality needs to be collected in at least two task rounds, corresponding data collection tasks can be set for each model in each task round. This forms the correspondence between modalities and tasks. Since tasks have an execution order, mapping data modalities to tasks allows for further consideration of the task execution order during the modeling process, making the modeling data more intuitive, regular, and easy to query and statistically analyze.
[0082] The resulting modeling data has a three-dimensional structure with site structure, task round structure, and processing stage as its three dimensions. Figure 3 This is a schematic diagram of the site structure provided by the present invention, as shown below. Figure 3 As shown, the site structure consists of a first level of sites and a second level of subjects under each site, forming a hierarchical structure. Figure 4 This is a schematic diagram of the task round structure provided by the present invention, as shown below. Figure 4 As shown, the modal data contained in a round also present a two-level hierarchical structure. In the task round structure, the first level is the modalities under the task round, and the second level is the task to which each modality belongs. Furthermore, there is a clear temporal sequence relationship between at least two task rounds, and the tasks within the same task round are arranged according to the temporal sequence of measurement and data collection. Figure 4 In this context, BEHV represents cognitive behavioral data, MRI represents magnetic resonance imaging data, EEG represents electroencephalography data, BIO represents biological sample data, and ENV represents environmental variable data.
[0083] In the method provided by this invention, the data structure of the modeling data is based on three dimensions: site structure, task round structure, and processing stage. The site structure and task round structure are both hierarchical structures, which can more comprehensively reflect the characteristics of the modeling data in each dimension, so that data users of different roles can quickly locate useful information from it.
[0084] Based on any of the above embodiments, the method further includes the following after step 120:
[0085] Based on the modeling data, queue data processing is performed.
[0086] Specifically, for the three-dimensional modeling data obtained from modeling, queue data processing can be performed. Here, queue data processing can be used to perform macro-level statistics on the queue data, such as counting the number of participants in a task, or counting the quality control progress of a task, or counting the distribution of participants who have completed at least two rounds. In addition, queue data processing can also use the modeling data as the basis for retrieval, and query data that meets the user's expectations from the modeling data according to the keywords entered by the user, so as to achieve rapid retrieval of queue data. This embodiment of the invention does not make specific limitations in this regard.
[0087] Based on any of the above embodiments Figure 5 This is a flowchart illustrating the modeling data processing method provided by the present invention, as shown below. Figure 5 As shown, step 120 is followed by:
[0088] Step 131: Determine the target parameters for the dimension to be statistically analyzed.
[0089] Here, the dimension to be counted can be any one of the three dimensions. The dimension to be counted varies depending on the statistical objective. For example, when counting the number of participants in a certain task, the dimension to be counted is the task. When analyzing the data of a participant under various tasks, the dimension to be counted is the participant.
[0090] The target parameters under the dimension to be analyzed are used to limit the scope of the statistical objective. The target parameters can be entered or specified by the statistician, or they can be automatically generated by the system. For example, when the dimension to be analyzed is a task, the target parameter can be all tasks in a certain round, or a specific task. When the dimension to be analyzed is a participant, the target parameter can be all participants at a certain site, or a specific participant.
[0091] Step 132: Based on the position of the target parameter in the dimension to be counted in the modeling data, determine the planar data that conforms to the target parameter from the modeling data.
[0092] Specifically, after determining the target parameter in the dimension to be statistically analyzed, the position of the target parameter can be located in the modeling data. Here, since the modeling data is a three-dimensional structure, after determining the position of the target parameter in one dimension, the planar data where the target parameter is located can be extracted from the modeling data based on the position of the target parameter.
[0093] Step 133: Perform data statistics based on the planar data.
[0094] Specifically, after obtaining the planar data of the target parameter, data statistics can be performed on the planar data to obtain the statistical results of the corresponding statistical target.
[0095] Based on any of the above embodiments, when the dimension to be counted is a task, the target parameter determined in step 131 can be one or more specific tasks, and the planar data obtained in step 132 is the data of each subject in each processing stage corresponding to one or more specific tasks.
[0096] Accordingly, step 133 may include:
[0097] Based on the projection length of the planar data in the direction corresponding to the subject, the amount of data for the task corresponding to the target parameter is determined.
[0098] Specifically, the planar data here can be data constructed by taking the subjects and the treatment phase as two coordinate axes on a plane. By calculating the projection length of the planar data in the direction corresponding to the subjects, that is, the projection length of the planar data on the coordinate axis to which the subjects belong, the number of subjects participating in the task indicated by the target parameter can be determined, that is, the amount of data for the task corresponding to the target parameter.
[0099] It should be noted that when there are at least two target parameters, the projection length of the planar data corresponding to each target parameter in the corresponding direction of the subject can be calculated separately, and the sum of the projection lengths of each target parameter can be used as the amount of data for all target parameters corresponding to the task.
[0100] For example, based on Figure 2 The modeling data structure shown can determine all planes parallel to the subject-processing coordinate axis plane in the modeling data based on the position of the target parameters on the task coordinate axis, thus obtaining planar data. Based on this, the projection length of each planar data point on the subject coordinate axis is calculated to obtain the data volume for each task indicated by the target parameters. Summing up all projection lengths yields the data volume for all tasks corresponding to the target parameters.
[0101] In addition, step 133 may also include:
[0102] The data in the planar data whose processing stage is the data acquisition stage is identified as acquired data, and the data in the planar data whose processing stage is the quality control stage is identified as quality control data.
[0103] Based on the projection length of the collected data in the direction corresponding to the subject, and the projection length of the quality control data in the direction corresponding to the subject, the quality control progress of the task corresponding to the target parameter is determined.
[0104] Specifically, the planar data here can be data constructed using the subjects and treatment stages as two coordinate axes on a plane. In the planar data, with the treatment stage as the statistical dimension, data corresponding to each treatment stage can be obtained, including the data collection data corresponding to the data collection stage and the quality control data corresponding to the quality control stage. Among them, the data collection data includes the data of each subject in the data collection stage under the task corresponding to the target parameter, and the quality control data includes the data of each subject in the quality control stage under the task corresponding to the target parameter.
[0105] By statistically analyzing the projection length of the collected data onto the subject's corresponding direction, the amount of data in the data collection phase of the task corresponding to the target parameters can be obtained. Similarly, by statistically analyzing the projection length of the quality control data onto the subject's corresponding direction, the amount of data in the quality control phase of the task corresponding to the target parameters can be obtained. By calculating the ratio between the amount of data in the quality control phase and the amount of data in the data collection phase, the quality control progress of the task corresponding to the target parameters can be determined.
[0106] Furthermore, the quality control stage can be positioned at the site quality control stage to determine the site quality control progress of the task corresponding to the target parameters; alternatively, the quality control stage can be positioned at the center quality control stage to determine the center quality control progress of the task corresponding to the target parameters.
[0107] For example, based on Figure 2 The modeling data structure shown allows us to determine all planes parallel to the subject-processing axis plane in the modeling data, based on the position of the target parameters on the task coordinate axis; this results in planar data. On this basis, we calculate the projection length X of the planar data on the subject coordinate axis during the processing stage (data acquisition stage), and sum the projection lengths during all processing stages (data acquisition stage) to obtain ∑X. Furthermore, we calculate the projection length X′ of the planar data on the subject axis during the processing stage (site quality control stage), and sum the projection lengths during all processing stages (site quality control stage) to obtain ∑X′. Calculating η = ∑X′ / ∑X yields the site quality control progress.
[0108] Based on any of the above embodiments, when the dimension to be statistically analyzed is the subject, the target parameter determined in step 131 can be one or more specific subjects, or the site where the subject is located. The planar data obtained in step 132 is the data of each task participated in by one or more specific subjects at each processing stage.
[0109] Accordingly, step 133 includes:
[0110] Determine the current task round;
[0111] Based on the projection length of the planar data in the direction corresponding to the task, the task round of the latest task corresponding to the target parameter is determined;
[0112] Remove the target parameters of the task rounds in which the latest task is located before the current round, and perform subject statistics based on the remaining target parameters.
[0113] Specifically, the recruitment and distribution of participants across multiple sites are crucial for guiding participant replenishment during cohort construction and maintenance. When conducting participant statistics, the target parameter is the range of participants whose recruitment and distribution statistics need to be compiled; for example, it could be all participants across multiple sites. Furthermore, it is necessary to determine the current round of the longitudinal cohort study, i.e., the current task round.
[0114] By statistically analyzing the projection lengths of the planar data corresponding to each participant as indicated in the target parameters onto the corresponding task direction, the task round in which each participant participated in the latest task, as indicated in the target parameters, is obtained. By comparing the current task round with the task round in which each participant participated in the latest task, as indicated in the target parameters, participants who did not participate in the current task round can be identified. These participants, whose latest task round is earlier than the current task round, are then deleted, and only those who participated in the current task round are retained for participant statistics.
[0115] Furthermore, the participant statistics here can be specifically clustered at the station level, counting the number of planes belonging to that station and drawing a histogram; or the participant distribution can be statistically analyzed according to demographic data such as age and gender, specifically by using age or gender as the granularity to analyze the selected target parameters.
[0116] For example, based on Figure 2 The modeling data structure shown can determine the current task round wl, and based on the position of the target parameter on the subject's coordinate axis, identify all planes in the modeling data that are parallel to the task-processing phase coordinate axis plane, thus obtaining planar data. Based on this, the projection length of the planar data onto the task coordinate axis is calculated, which represents the latest task, thereby determining the task round w of the latest task. j If w j <w l If the condition is met, the corresponding planar data is discarded; otherwise, it is retained, until the comparison operation of the latest task round corresponding to all planar data is completed. For the retained planar data, participant statistics are performed.
[0117] Currently, information technology tools serving cohort research construction primarily focus on specific scenarios (such as data collection, quality control, and analysis). Cohort builders mainly rely on searching relevant papers or official websites at different stages of the project to obtain relevant work experience and plan the overall technical path. However, the data, tools, and standards required for the research are all independently developed standards for each cohort data information plane. This fragmented information hinders timely and effective reference for cohort builders. Furthermore, the lack of a unified data modeling framework prevents these information technology tools from being effectively integrated according to a unified standard, thus hindering their large-scale adoption in cohort research management.
[0118] Modeling queue data within the business context of longitudinal cohort studies will help provide a unified framework for dynamic, multidimensional data in longitudinal cohort studies and establish a unified understanding of the data among different roles in cohort studies. This will play a crucial role in standardizing the standard operating procedures for queue data management and further improving data quality.
[0119] To address this issue, based on any of the above embodiments, step 120 is followed by:
[0120] Receive query information;
[0121] Based on the retrieval platform, the query information is retrieved to obtain the retrieval results;
[0122] The retrieval platform is built based on the modeling data, as well as information toolsets and / or project information, which are stored in a structure that includes at least two dimensions: tasks and processing phases.
[0123] Specifically, in response to the problem of scattered information related to longitudinal cohort research in related technologies, this invention provides a retrieval platform that combines modeling data with information toolsets and / or project information to facilitate the integration of various types of information required for longitudinal cohort research, thereby improving information retrieval efficiency and promoting the development of longitudinal cohort research.
[0124] The retrieval platform here not only stores modeling data obtained from data modeling using subjects, tasks, and treatment phases as three dimensions, but also stores information toolsets and / or project information required for longitudinal cohort studies. To facilitate information aggregation and retrieval, information toolsets and / or project information can be stored according to a dimensional structure that includes at least tasks and treatment phases.
[0125] Furthermore, specifically during storage, for information toolsets and / or project information, information can be collected according to the table below and stored in conjunction with the tasks and processing stages. This information collection can be achieved using Web of Science, Scopus, PubMed, etc.
[0126]
[0127] When conducting information retrieval based on the aforementioned retrieval platform, the received query information can be keywords required for information retrieval. The setting of these keywords can refer to one or more dimensions of the modeling data, namely, subjects, tasks, and processing stages. The retrieval results obtained can include modeling data (Database) related to the query information, or information toolsets (Toolkit) and / or project information related to the query information. This embodiment of the invention does not specifically limit this.
[0128] Based on any of the above embodiments, the processing stage includes a data acquisition stage, a quality control stage, an analysis stage, and a sharing stage.
[0129] The quality control stage includes a site quality control stage and a central quality control stage. The site quality control stage is used to perform quality control on data collected from the same site, and the central quality control stage is used to perform quality control on data collected from at least two sites.
[0130] Specifically, during the data aggregation process, to improve the success rate of a single data collection, quality control of the data needs to be conducted on-site at the subject collection location. Simultaneously, considering the varying data quality control standards or specific implementation measures at different sites, the center can also re-check the data quality of each site and provide feedback shortly after data collection. In this process, the data quality control performed at the site is called site quality control, and the data quality control performed at the center is called central quality control. Central quality control occurs after site quality control is passed. Specifically, each site can perform site quality control on the queue data it collects. Site quality control can be completed within 10 minutes or half an hour after data collection is completed. Furthermore, each site can send the queue data that has passed site quality control to the center at regular intervals for further quality control. Data transmission between sites and the center can be achieved through the local server at each site. Central quality control can be completed within half a day or a day after data collection is completed.
[0131] The method provided in this invention, by setting up two levels of quality control, ensures that queue data collected from multiple sites can maintain a consistent quality standard, which helps to improve the overall quality of queue data.
[0132] Based on any of the above embodiments, when studying the brain development of a certain type of children, multimodal data of that type of children can be collected. Figure 6 This is a schematic diagram of the structure of the multimodal data provided by the present invention, as shown below. Figure 6 As shown, multimodal data can be categorized into four types based on data modality: cognitive psychological and behavioral data, magnetic resonance imaging (MRI) data, electroencephalogram (EEG) data, and biological sample data. The task classification for each modality can be represented as follows: Figure 6 The next level of the intermediate modal data.
[0133] After data collection is completed, data quality control can be performed at both the site and center levels. The data after quality control is then cleaned. During data cleaning, structured and unstructured data can be cleaned separately.
[0134] Cleaning of structured data: For structured data, real-time desensitization is mainly performed during the query and retrieval of sensitive data. The sensitive and privacy information involved includes names, ID numbers, and mobile phone numbers. Desensitization can be achieved through various common desensitization techniques, such as SQL (Structured Query Language) statement rewriting technology, which performs function operations on sensitive fields so that the real data can be retrieved or exported after desensitization.
[0135] For unstructured data, such as MRI, EEG, and audio data, different methods can be used for desensitization. For example, for audio data, differential privacy technology can be used to incorporate customized noise to desensitize images and audio signals. For MRI data, image analysis and synthesis techniques can also be used to achieve desensitization, such as using Gaussian noise "mosaic" to remove facial information to protect the privacy of image data.
[0136] Following the steps above, the queue data can be obtained. This queue data involves at least two modalities from multiple sources, including cognitive-psychological-behavioral data, magnetic resonance imaging (MRI), electroencephalography (EEG), biological samples, and the environment. Given a set of multimodal data types:
[0137] M={BEHV, MRI, EEG, BIO, ENV}
[0138] Among them, BEHV represents cognitive behavioral data, MRI represents magnetic resonance imaging (MRI) data, EEG represents electroencephalogram (EEG) data, BIO represents biological samples data, and ENV represents environmental data.
[0139] Each data modality consists of several tasks, and the tasks included in each data modality are as follows:
[0140] BEHV={behv1, behv2,..., behv N}
[0141] MRI={mri1, mri2, ...., mri N′}
[0142] EEG={eeg1, eeg2, ...., eeg N″}
[0143] BIO={bio1, bio2,..., bio N″′}
[0144] ENV={env1,env2,...,env N″″}
[0145] In the above set, behv1, mri1, eeg1, bio1, and env1 are all specific tasks, and N, N′, N″, N″′, and N″″ are not necessarily equal, and are often not equal.
[0146] Given a set of multiple sites in cohort studies:
[0147]
[0148] Where K is the total number of stations and K ≠ 0, s k This is one of the stations. During the actual cohort study, station s... k It mainly includes two types of personnel: subjects and experimenters.
[0149] The subject (or participant) in a cohort study refers to the individuals who contribute and collect data; the experimenter (or researcher) is the person who conducts the cohort study, responsible for data collection, processing, and other related procedures. k The definitions of experimenter and subject are as follows:
[0150] sub k ={sub k1 , ...,sub ki , ...,sub kn}
[0151] exp k ={exp k1 , ...,exp kj , ...,exp kn}
[0152] Where sub ki For site s k The i-th subject in the experiment, exp kj For site s k The j-th examiner, sub ki ∈s k ,exp kj ∈s k .
[0153] Considering the characteristics of the entire process in cohort research, given the enumeration set of process stages:
[0154]
[0155] Ph1 to Ph5 correspond to five processes: data collection, local quality control, central quality control, analysis, and sharing, respectively. Participation, but At pH 1, participate.
[0156] Given a set of rounds in a cohort study, and considering the task rounds involved:
[0157]
[0158] Among them, w l That is, the first task round, the task round satisfies the following conditions:
[0159]
[0160] when This type of study belongs to the longitudinal cohort study, when This type of study belongs to cross-sectional studies;
[0161]
[0162] End time.
[0163] Based on the definitions of the basic concepts in cohort studies mentioned above, a formal definition can be made of cohort studies, the data generated by the cohort, and their related constraints:
[0164] This refers to a cohort study. A cohort study is a type of study involving multiple sites. Collect multimodal data And targeting sub ki Multiple rounds Follow-up and tracking, and full-process analysis of the generated data. This is a type of research that involves processing and analyzing data. Therefore, the quadruple definition of a cohort study can be derived:
[0165]
[0166] in Indicates the length of the vector.
[0167] Regarding the data generation process for the queue, the generation of queue data occurs in ph1, and the process is initiated by the experimenter. kj Select a specific task (such as behv) i ), targeting the subject sub ki At a certain site s k In a certain round w l The event is carried out at a certain time t, and the data it generates contains the above information, namely:
[0168]
[0169] in This indicates the collection of data for a specific task within a specific modality, i.e. It can be or or or or wait.
[0170] For the cohort data represented in this way, subjects, tasks, and processing stages can be used as the three dimensions of a cube model to model the cohort data, thereby obtaining modeling data that can also be represented from the three dimensions of subjects, tasks, and processing stages.
[0171] In the modeling data, due to the different computational indicators and measures for different modalities, the unit cube model is presented in the form of an object. When the task coordinate axis of the cube model is mapped to a cognitive psychological behavior modality, the data in this object takes on a structured form, which can be implemented in XML, JSON, or other storage formats. When the task coordinate axis of the cube model is mapped to a magnetic resonance imaging (MRI) sequence task, the data in this object includes image data formats such as DICOM and NifTI. Meanwhile, the various dimensions on the subject, task, and treatment phase coordinate axes still exhibit hierarchical structures or time-series relationships.
[0172] Based on any of the above embodiments, the construction process of a longitudinal cohort study can be abstracted into eight major scenarios, including multi-site subject sampling, large-scale subject tracking, dynamic management of test administration tasks, multimodal data acquisition, quality control and preprocessing, storage and management, computation and analysis, and openness and sharing.
[0173] Figure 7 This is a schematic diagram of a vertical queue scenario provided by the present invention, such as... Figure 7As shown, the above business scenarios can be mapped to the three-dimensional structure of the modeling data. For example, multi-site subject sampling and large-scale subject tracking correspond to the subject coordinate axis in the cube model, focusing on acquiring and maintaining subjects; dynamic management of administration tasks corresponds to the task coordinate axis in the cube model, with the main purpose of constructing a two-level structure of "modality-task / sequence" and dynamically maintaining each round w. l The testing tasks for each modality; in addition, the process from acquiring multimodal data to open sharing corresponds to the application stage coordinate axis in the cube model, which is used to identify the five processes of data from entry to exit.
[0174] also, Figure 7 The system involves eight modules: infrastructure, data aggregation, data storage, computing and analysis, open sharing, application systems, operation and maintenance, and security. These are all modules included in the retrieval platform. The connections between these modules and scenarios are used to illustrate the relationship between modules and scenarios.
[0175] The method provided in this invention, based on the research content of longitudinal cohort studies, outlines the business characteristics of cohort studies and the multimodal data task system. A two-level quality control system ensures the validity and completeness of the collected multimodal data. Furthermore, unlike traditional modeling methods that use indicators and their measurements as basic elements, this invention provides a method that uses the data objects generated from each interaction between the subject and the testing task as the basic granularity. It constructs a cube model with three dimensions: Subject, Task, and Phase, and uses the data objects as the unit cubes within this model.
[0176] Based on this, the embodiments of the present invention also establish a unified understanding of the data among multiple parties ("subjects" - data generators, "experimenters" - data maintainers, and "researchers" - data users) according to the cube model. Under this framework, the data is mapped to the main modules of the data information platform for cohort research, and a retrieval platform for cohort research is established, thereby ensuring the efficiency and scientific rigor of cohort research.
[0177] Based on any of the above embodiments Figure 8 This is a schematic diagram of the structure of the queue data processing device provided by the present invention, as shown below. Figure 8 As shown, the device includes:
[0178] The data determination unit 810 is used to generate queue data based on multimodal data. The multimodal data is obtained by collecting data on tasks in at least two task rounds at at least two stations during the queue study. The multimodal data includes data from at least two modalities, and each station includes at least two subjects.
[0179] The data modeling unit 820 is used to perform data modeling on the queue data based on the cube model as the basic element to obtain modeling data. The cube model is a three-dimensional structure with subjects, tasks and processing stages as the three dimensions.
[0180] The apparatus provided in this invention is based on a cube model with subjects, tasks, and processing stages as three dimensions. It performs data modeling on the data obtained from cohort research and processes the modeled data. This effectively avoids the problem of wasted storage space, solves the business needs of macroscopic statistics of multimodal data in longitudinal cohort research, and helps to improve the efficiency of cohort data processing.
[0181] Based on any of the above embodiments, the data determination unit is used for:
[0182] Based on the tasks corresponding to the data of the at least two modalities and the task rounds to which the tasks belong, and / or
[0183] The data from the at least two modalities, along with the subjects and their respective stations, are integrated to obtain the queue data.
[0184] Based on any of the above embodiments, the data modeling unit is used for:
[0185] Using the cube model as the basic element, and based on the tasks and task rounds corresponding to the modalities of each basic data in the queue data, as well as the subjects and processing stages corresponding to each basic data, modeling is performed to obtain the modeling data.
[0186] The data structure of the modeling data is a three-dimensional structure with three dimensions: site structure, task round structure, and processing stage. The site structure is a hierarchical structure including sites and subjects under each site, and the task round structure is a hierarchical structure including task rounds and tasks under each task round.
[0187] Based on any of the above embodiments, the device further includes a first processing unit, configured to:
[0188] Determine the target parameters for the dimension to be statistically analyzed;
[0189] Based on the position of the target parameter in the dimension to be counted in the modeling data, planar data that conforms to the target parameter is determined from the modeling data;
[0190] Based on the aforementioned planar data, statistical analysis is performed.
[0191] Based on any of the above embodiments, when the dimension to be counted is a task, the first processing unit is used to:
[0192] Based on the projection length of the planar data in the direction corresponding to the subject, the amount of data for the task corresponding to the target parameter is determined.
[0193] And / or,
[0194] The data in the planar data whose processing stage is the data acquisition stage is identified as acquired data, and the data in the planar data whose processing stage is the quality control stage is identified as quality control data.
[0195] Based on the projection length of the collected data in the direction corresponding to the subject, and the projection length of the quality control data in the direction corresponding to the subject, the quality control progress of the task corresponding to the target parameter is determined.
[0196] Based on any of the above embodiments, when the dimension to be statistically analyzed is the subject, the first processing unit is used to:
[0197] Determine the current task round;
[0198] Based on the projection length of the planar data in the direction corresponding to the task, the task round of the latest task corresponding to the target parameter is determined;
[0199] Remove the target parameters of the task rounds in which the latest task is located before the current round, and perform subject statistics based on the remaining target parameters.
[0200] Based on any of the above embodiments, the device further includes a second processing unit, used for:
[0201] Receive query information;
[0202] Based on the retrieval platform, the query information is retrieved to obtain the retrieval results;
[0203] The retrieval platform is built based on the modeling data, as well as information toolsets and / or project information, which are stored in a structure that includes at least two dimensions: tasks and processing phases.
[0204] Based on any of the above embodiments, the quality control stage includes a site quality control stage and a central quality control stage. The site quality control stage is used to perform quality control on data collected from the same site, and the central quality control stage is used to perform quality control on data collected from at least two sites.
[0205] Figure 9 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 9As shown, the electronic device may include: a processor 910, a communication interface 920, a memory 930, and a communication bus 940, wherein the processor 910, the communication interface 920, and the memory 930 communicate with each other through the communication bus 940. The processor 910 can call logical instructions in the memory 930 to execute a data processing method, which includes:
[0206] Based on multimodal data, cohort data is generated. The multimodal data is obtained by collecting data on tasks in at least two task rounds at at least two stations during the cohort study. The multimodal data includes data from at least two modalities, and each station includes at least two subjects.
[0207] Using a cube model as the basic element, data modeling is performed on the queue data to obtain modeling data. The cube model is a three-dimensional structure with subjects, tasks, and processing stages as the three dimensions.
[0208] Furthermore, the logical instructions in the aforementioned memory 930 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0209] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program, the computer program being able to be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer being able to execute the data processing methods provided by the above methods, the method comprising:
[0210] Based on multimodal data, cohort data is generated. The multimodal data is obtained by collecting data on tasks in at least two task rounds at at least two stations during the cohort study. The multimodal data includes data from at least two modalities, and each station includes at least two subjects.
[0211] Using a cube model as the basic element, data modeling is performed on the queue data to obtain modeling data. The cube model is a three-dimensional structure with subjects, tasks, and processing stages as the three dimensions.
[0212] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the data processing methods provided by the methods described above, the method comprising:
[0213] Based on multimodal data, cohort data is generated. The multimodal data is obtained by collecting data on tasks in at least two task rounds at at least two stations during the cohort study. The multimodal data includes data from at least two modalities, and each station includes at least two subjects.
[0214] Using a cube model as the basic element, data modeling is performed on the queue data to obtain modeling data. The cube model is a three-dimensional structure with subjects, tasks, and processing stages as the three dimensions.
[0215] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0216] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0217] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data processing method, characterized in that, include: Based on multimodal data, queue data is generated. The multimodal data is obtained by collecting data on tasks under at least two task rounds at at least two stations during the queue study. The multimodal data includes data from at least two modalities, and each station includes at least two subjects. When generating the queue data based on the multimodal data, the data of each modality obtained under each task round is mapped to the task under each task round. Using a cube model as the basic element, data modeling is performed on the queue data to obtain modeling data. The cube model is a three-dimensional structure with subjects, tasks, and processing stages as the three dimensions. After obtaining the modeling data, the process also includes: In the modeling data, determine the planar data corresponding to the target parameter under the dimension to be statistically analyzed; When the dimension to be statistically analyzed is a task, the amount of data corresponding to the target parameter is determined based on the projection length of the planar data in the direction corresponding to the subject. The target parameter is one or more specific tasks, and the planar data is the data of each subject in each processing stage corresponding to one or more specific tasks. as well as, The data in the planar data whose processing stage is the data acquisition stage is identified as acquired data, and the data in the planar data whose processing stage is the quality control stage is identified as quality control data. Based on the projection length of the collected data in the direction corresponding to the subject, and the projection length of the quality control data in the direction corresponding to the subject, the quality control progress of the task corresponding to the target parameter is determined.
2. The data processing method according to claim 1, characterized in that, The generation of queue data based on multimodal data includes: Based on the tasks corresponding to the data of the at least two modalities and the task rounds to which the tasks belong, and / or the subjects corresponding to the data of the at least two modalities and the stations to which the subjects belong, the at least two modal data are integrated to obtain the queue data.
3. The data processing method according to claim 1, characterized in that, The process of modeling the queue data using a cube model as the basic element to obtain modeling data includes: Using the cube model as the basic element, and based on the tasks and task rounds corresponding to the modalities of each basic data in the queue data, and / or the subjects and processing stages corresponding to each basic data, modeling is performed to obtain the modeling data; The data structure of the modeling data is a three-dimensional structure with three dimensions: site structure, task round structure, and processing stage. The site structure is a hierarchical structure including sites and subjects under each site, and the task round structure is a hierarchical structure including task rounds and tasks under each task round.
4. The data processing method according to claim 1, characterized in that, The step of determining the planar data corresponding to the target parameter under the dimension to be statistically analyzed in the modeling data includes: Determine the target parameters under the dimension to be statistically analyzed; Based on the position of the target parameter in the dimension to be counted in the modeling data, planar data that conforms to the target parameter is determined from the modeling data.
5. The data processing method according to claim 1, characterized in that, After determining the planar data corresponding to the target parameter under the dimension to be statistically analyzed in the modeling data, the method further includes: Given that the dimension to be statistically analyzed is the subject, determine the current task round; Based on the projection length of the planar data in the direction corresponding to the task, the task round of the latest task corresponding to the target parameter is determined; Remove the target parameters of the task rounds in which the latest task is located that are earlier than the current round, and perform subject statistics based on the remaining target parameters.
6. The data processing method according to claim 1, characterized in that, The process of using a cube model as the basic element to model the queue data, obtaining modeled data, further includes: Receive query information; Based on the retrieval platform, the query information is retrieved to obtain the retrieval results; The retrieval platform is built based on the modeling data, as well as information toolsets and / or project information, which are stored in a structure that includes at least two dimensions: tasks and processing phases.
7. The data processing method according to any one of claims 1 to 6, characterized in that, The at least two modalities include at least two from cognitive psychological behavior, magnetic resonance imaging, electroencephalography, biological samples, and the environment; The processing phase includes at least two of the following: data acquisition phase, quality control phase, analysis phase, and sharing phase.
8. The data processing method according to claim 7, characterized in that, The quality control phase includes a site quality control phase and a central quality control phase. The site quality control phase is used to perform quality control on data collected from the same site, and the central quality control phase is used to perform quality control on data collected from at least two sites.
9. A data processing apparatus, characterized in that, include: A data determination unit is used to generate queue data based on multimodal data. The multimodal data is obtained by collecting data from at least two stations for at least two task rounds during the queue study. The multimodal data includes data from at least two modalities, and each station includes at least two subjects. When generating the queue data based on the multimodal data, the data of each modality obtained in each task round is mapped to the task in each task round. The data modeling unit is used to perform data modeling on the queue data based on the cube model as the basic element to obtain modeling data. The cube model is a three-dimensional structure with subjects, tasks and processing stages as the three dimensions. After obtaining the modeling data, the process also includes: In the modeling data, determine the planar data corresponding to the target parameter under the dimension to be statistically analyzed; When the dimension to be statistically analyzed is a task, the amount of data corresponding to the target parameter is determined based on the projection length of the planar data in the direction corresponding to the subject. The target parameter is one or more specific tasks, and the planar data is the data of each subject in each processing stage corresponding to one or more specific tasks. as well as, The data in the planar data whose processing stage is the data acquisition stage is identified as acquired data, and the data in the planar data whose processing stage is the quality control stage is identified as quality control data. Based on the projection length of the collected data in the direction corresponding to the subject, and the projection length of the quality control data in the direction corresponding to the subject, the quality control progress of the task corresponding to the target parameter is determined.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the data processing method as described in any one of claims 1 to 8.
11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the data processing method as described in any one of claims 1 to 8.