An arthritis data processing method and system based on ceRNA-network analysis

By identifying transcriptional consistency and targeted co-annotation relationships in arthritis samples, combined with pathological pathway enrichment verification, and calculating the regulatory influence of ceRNA network nodes, the limitations of traditional methods in arthritis data processing are overcome, enabling accurate identification of arthritis-specific core hub nodes and effective screening of biomarkers.

CN122369607APending Publication Date: 2026-07-10THE SECOND AFFILIATED HOSPITAL OF SHANDONG UNIV OF TRADITIONAL CHINESE MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THE SECOND AFFILIATED HOSPITAL OF SHANDONG UNIV OF TRADITIONAL CHINESE MEDICINE
Filing Date
2026-06-03
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Traditional ceRNA network analysis-based methods for processing arthritis data struggle to fully capture complex transcriptional associations and target co-annotation relationships, leading to limitations in the screening process, failing to accurately reflect the molecular mechanisms of arthritis, and affecting the screening of key RNA molecules and the identification of potential biomarkers.

Method used

By acquiring mRNA and non-coding RNA transcriptional profiles of arthritis samples, we identified transcriptional consistency patterns and target co-annotation relationships, generated an exclusion set for joint screening of transcriptional targets in arthritis, and combined pathological pathway enrichment and functional matching verification to calculate regulatory priorities and the regulatory influence of ceRNA network nodes, thus identifying arthritis-specific core hub nodes.

Benefits of technology

It improved the accuracy of arthritis data analysis, optimized the screening process for potential biomarkers, ensured the biological rationality and reliability of the screening results, and enhanced the ability to identify key disease-related nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122369607A_ABST
    Figure CN122369607A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of biological information data processing, in particular to a kind of arthritis data processing method and system based on ceRNA-network analysis, comprising the following steps: obtaining arthritis sample RNA data and target annotation, extract shared element screening regulation pair, by pathological pathway enrichment and function verification, measure correlation to determine priority, integrate ceRNA network structure to evaluate regulatory influence, parse node role to adjust centrality, generate specific core hub node identification atlas.In the present application, by introducing dynamic transcription consistency pattern recognition and transcription correlation and target co-annotation relationship joint comparison, the accuracy of arthritis data analysis is improved, the evaluation of candidate regulatory pair is strengthened, the result reliability and regulatory confidence level are improved, combined with ceRNA network node regulatory influence evaluation, the identification ability of key node is enhanced, effectively solve the blind spot of traditional method, ensure the accurate identification of arthritis specific core hub node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics data processing technology, and in particular to a method and system for processing arthritis data based on ceRNA-network analysis. Background Technology

[0002] The field of bioinformatics data processing technology primarily involves using interdisciplinary methods such as computer science, statistics, and biology to process and analyze large amounts of biological data, aiming to extract biologically significant information. This field encompasses the processing of various biological data types, including RNA molecular omics, transcriptomics, and proteomics, encompassing RNA transcription data analysis, protein-protein interaction network construction, and RNA mutation and disease association analysis. Core technologies include data mining, machine learning, and network biology methods, used to analyze complex molecular mechanisms within organisms and their relationship with diseases. With the development of high-throughput sequencing technology, bioinformatics data processing methods are becoming increasingly precise, enabling the identification of potential disease biomarkers from massive amounts of biological data, thus driving the development of precision medicine.

[0003] Traditional ceRNA-network analysis-based methods for arthritis data processing involve constructing ceRNA network models and combining them with high-throughput sequencing data to explore the regulatory relationships between non-coding RNAs and mRNAs. In research on diseases such as arthritis, this method is used to screen key biomarkers and reveal their roles in disease development. Traditional ceRNA network analysis methods typically rely on known mRNA and non-coding RNA transcriptional data, using bioinformatics methods to calculate their interactions and infer their potential mechanisms in the occurrence and development of arthritis. Specific analytical steps include data preprocessing, RNA transcriptional differential analysis, ceRNA network construction, and screening of key RNA molecules. The aim is to discover disease-related biomarkers by elucidating the complex regulatory relationships between molecules, providing a theoretical basis for the diagnosis and treatment of arthritis.

[0004] Traditional ceRNA network analysis-based methods for processing arthritis data rely on known mRNA and non-coding RNA transcriptional data, making it difficult to fully capture complex transcriptional associations and target co-annotation relationships, thus limiting their screening process. These methods largely depend on static data analysis, failing to dynamically adjust and consider transcriptional consistency patterns across different samples. This results in regulatory networks that cannot accurately reflect the actual molecular mechanisms of diseases such as arthritis. Furthermore, traditional methods neglect the deep correlation between RNA molecule functional annotation and signaling pathway enrichment, leading to a failure to effectively integrate pathological pathways during the screening of key RNA molecules. This hinders a comprehensive assessment of the biological function and clinical relevance of candidate regulatory pairs, affecting the accurate identification of potential biomarkers and ultimately limiting their practical application in the diagnosis and treatment of arthritis. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the existing technology and to propose a method and system for processing arthritis data based on ceRNA-network analysis.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: a method for processing arthritis data based on ceRNA-network analysis, comprising the following steps: S1: Obtain transcriptional profile data of mRNA and non-coding RNA, trace RNA target annotation information and sample grouping labels from arthritis samples, extract candidate regulatory pairs with shared trace RNA response elements, simultaneously identify transcriptional consistency patterns between samples, jointly compare transcriptional association and target co-annotation relationship, perform regulatory association labeling, and generate a joint screening and exclusion set of transcriptional targets for arthritis. S2: Based on the arthritis transcription-targeted joint screening regulatory exclusion set, extract RNA molecular function annotation fields and pathway enrichment fields, match molecular function tags and signaling pathway tags, and generate a pathological pathway enrichment and functional matching verification identifier set. S3: Obtain candidate regulatory pairs corresponding to the pathological pathway enrichment and functional matching verification identifier set, unify the transcription unit, classify the association strength interval and regulatory confidence level, and obtain the regulatory priority determination result. S4: Based on the regulatory priority determination results, identify the set of connection edges of candidate regulatory pairs in the ceRNA network, extract the number of directly adjacent nodes and the number of indirect path nodes, integrate multi-level neighborhood structure information, inject regulatory confidence level parameters, and obtain the ceRNA network node regulatory impact assessment dataset.

[0007] As a further embodiment of the present invention, the arthritis transcription-targeted joint screening and regulatory exclusion set includes regulatory failure status, exclusion triggering basis, and exclusion range; the pathological pathway enrichment and functional matching verification identifier set includes pathway enrichment category, regulatory direction identifier, and activation confidence number; the regulatory priority determination result includes association strength type, functional consistency level, and regulatory intervention level; and the ceRNA network node regulatory impact assessment dataset includes direct connection weight, indirect path weight, and regulatory confidence adjustment weight.

[0008] As a further aspect of the present invention, the step of generating the arthritis transcription-targeted joint screening regulatory exclusion set specifically includes: S111: Obtain transcriptional profile data of mRNA and non-coding RNA, trace RNA targeted annotation information and sample grouping labels from arthritis samples. Extract the annotation terms belonging to the current regulatory pair according to the mapping relationship between regulatory pairs and trace RNA response elements, and establish a regulatory annotation set. S112: Call the regulatory annotation set, perform cross-sample consistency test on the transcriptional profile data of mRNA and non-coding RNA, extract transcriptional trend symbols and target co-annotation status, set co-transcription threshold, perform logical AND operation, and generate regulatory feasibility judgment result; S113: Extract unique identifiers for candidate regulatory pairs, identify whether the regulatory pairs lack common trace RNA targeting support based on the regulatory feasibility determination results, set regulatory failure status markers for regulatory pairs lacking support, and obtain the joint screening regulatory exclusion set for arthritis transcriptional targeting.

[0009] As a further aspect of the present invention, the steps for generating the pathological pathway enrichment and functional matching verification identifier set are as follows: S211: Identify the recorded information in the exclusion set of the transcriptional targeting joint screening regulation of arthritis, screen candidate regulatory pairs that are not marked as excluded, and extract the set of activatable regulatory pairs according to the mapping relationship between the regulatory pairs and the arthritis-related pathways. S212: Obtain the functional annotation fields and pathway enrichment fields of the corresponding mRNA and non-coding RNA of the set of activatable regulatory pairs, and generate a pathway association parameter set based on the path association between the RNA molecular ontology hierarchy and KEGG pathway number; S213: Call the pathway association parameter set, extract the molecular function category in the function annotation field and the arthritis-specific pathway identifier in the pathway enrichment field, calculate the pathway enrichment by combining the hypergeometric distribution test algorithm, determine whether the regulatory pair meets the enrichment condition that the P value is less than the preset threshold, and generate a pathological pathway enrichment and function matching verification identifier set.

[0010] As a further aspect of the present invention, the step of obtaining the control priority determination result specifically includes: S311: Obtain candidate regulatory pairs corresponding to the pathological pathway enrichment and functional matching verification identifier set, match the original transcriptome records according to the unique identifiers in the pathological pathway enrichment and functional matching verification identifier set, screen the corresponding regulatory entries, extract the regulatory pair identifier field, locate the matching record entries, screen the target regulatory pair numbers, and generate the target regulatory pair index set. S312: Call the target regulatory pair index set, extract the transcription vectors of mRNA and non-coding RNA in each corresponding target regulatory pair, and simultaneously identify the sample grouping field, calculate the Pearson correlation coefficient between the two transcription vectors, and generate a set of correlation coefficients. S313: Compare each coefficient in the set of correlation coefficients with the preset correlation strength range to determine whether the correlation strength of the target control pair is within the preset range, and generate a control priority determination result.

[0011] As a further aspect of the present invention, the steps for obtaining the dataset for assessing the regulatory impact of ceRNA network nodes specifically include: S411: Based on the regulatory priority determination result, analyze the adjacency relationship of the target regulatory pair node in the ceRNA network, locate its position in the global network according to the node's unique identifier, and extract the list of directly adjacent nodes and the list of two-hop adjacent nodes to generate a neighborhood structure data set. S412: Obtain the neighborhood structure data group, identify and regulate the association confidence level, and perform weighted labeling processing on the number of direct adjacencies and the number of indirect paths based on it to generate a weighted neighborhood parameter group. S413: Call the weighted neighborhood parameter group, and perform a linear combination operation based on the combined distribution characteristics of direct adjacency weights and indirect path weights in the current weighted state, using the formula: : Calculate the regulatory influence value to obtain a dataset assessing the regulatory influence of ceRNA network nodes; in, Represents the influence value of regulation. Representing the The adjacency weight between each node and the target node; Representing the The weight of the direct adjacency path from each node to the target node; Representing the Adjustment factor for each node, Represents a regulatory factor. This represents the total number of associated nodes participating in the calculation within the ceRNA network.

[0012] As a further aspect of the present invention, the method further includes step S5: S5: Based on the ceRNA network node regulation impact assessment dataset, the network role type corresponding to the target regulatory pair molecule is matched by the node's unique identifier. The role field is parsed and loaded into the regulatory context. The ceRNA network node regulation impact assessment dataset and role type parameters are linked together. The network centrality calculation method and output format are adjusted to generate an arthritis-specific core hub node identification map. The arthritis-specific core hub node identification map includes role type identifiers, number of direct connections, number of indirect paths, and regulatory influence value.

[0013] As a further aspect of the present invention, the steps for obtaining the arthritis-specific core hub node identification map are as follows: S511: Based on the ceRNA network node regulatory influence assessment dataset, identify the correspondence between the node's unique identifier and the target regulatory pair molecule, filter the role type records pointed to by the current regulatory network, and generate a target role type set; S512: Based on the target role type set, and according to the alignment relationship between role type identifier and regulatory priority, determine the centrality computational requirement of the corresponding role type based on the ceRNA network node regulatory impact assessment dataset, and generate a centrality parameter allocation group; S513: Write the calibrated calculation parameters in the centrality parameter allocation group into the centrality algorithm of the corresponding role type, identify the output state after the centrality calculation adjustment, and output the arthritis-specific core hub node identification map.

[0014] The arthritis data processing system based on ceRNA-network analysis is used to execute the aforementioned arthritis data processing method based on ceRNA-network analysis. The system includes: The regulation exclusion identification module obtains the transcriptional profile data of mRNA and non-coding RNA, trace RNA targeting annotation information and sample grouping labels in arthritis samples, and jointly judges whether the current regulatory pair meets the ceRNA action conditions, and generates an arthritis transcriptional targeting joint screening regulation exclusion set. The regulation activation judgment module identifies candidate regulatory pairs that are not marked as excluded from the joint screening regulation exclusion set of the arthritis transcriptional target, extracts RNA molecular functional annotation fields and pathway enrichment fields, determines whether the conditions for arthritis-related pathway enrichment are met, and generates a pathological pathway enrichment and functional matching verification identifier set. The regulatory priority analysis module calculates the correlation between the transcriptional vectors of mRNA and non-coding RNA based on the candidate regulatory pairs corresponding to the pathological pathway enrichment and functional matching verification identifier set, determines the correlation strength of the regulatory pairs, and obtains the regulatory priority determination result. Based on the regulatory priority determination results, the regulatory influence calculation module identifies the adjacency relationship of candidate regulatory pairs in the ceRNA network, and combines the regulatory association confidence level as a regulation term to calculate the ceRNA network node regulatory influence assessment dataset. The key node identification module matches the network role type corresponding to the target regulatory pair molecules based on the node's unique identifier. It adjusts the execution parameters of the centrality algorithm according to the ceRNA network node regulation influence assessment dataset to generate an arthritis-specific core hub node identification map.

[0015] Compared with the prior art, the advantages and positive effects of the present invention are as follows: This invention effectively improves the accuracy of arthritis data analysis by introducing dynamic transcriptional consistency pattern recognition and joint comparison of transcriptional association and target co-annotation relationships. The combination of dynamic screening of regulatory pairs and pathological pathway enrichment verification not only strengthens the correlation between RNA molecular functional annotation and signaling pathway enrichment but also optimizes the screening process for potential biomarkers, effectively ensuring the biological rationality of the screening results. By calculating the Pearson correlation coefficient and classifying the correlation strength intervals, the evaluation of candidate regulatory pairs is enhanced, improving the reliability and regulatory confidence level of the results. At the same time, combining the node regulatory impact assessment of the ceRNA network enhances the ability to identify key disease-related nodes, effectively solving the blind spots of traditional methods in molecular mechanism analysis and ensuring the accurate identification of arthritis-specific core hub nodes. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the workflow of the present invention; Figure 2 This is a flowchart illustrating the process of obtaining the exclusion set for the joint screening of transcriptional targets for arthritis in this invention. Figure 3 This is a flowchart illustrating the acquisition of the pathological pathway enrichment and functional matching verification identifier set in this invention. Figure 4 This is a flowchart illustrating the process of obtaining the priority determination result in this invention. Figure 5 This is a flowchart illustrating the process of obtaining the dataset for assessing the regulatory impact of ceRNA network nodes in this invention. Figure 6 This is a flowchart illustrating the process of obtaining the arthritis-specific core hub node identification map in this invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0018] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0019] All user-related information involved in this invention (including but not limited to biometric information, identity verification information, behavioral data, device information, and other data that can be used for identity verification and personalized services) is collected and processed with the user's full knowledge and voluntary consent. The collection, storage, and use of all information strictly comply with applicable national and regional laws and regulations, and meet relevant data protection standards and policy requirements. The use of data is limited to purposes necessary for providing the technical services of this invention, and reasonable technical and management measures will be taken to ensure the security and confidentiality of users' personal information in terms of information protection and privacy.

[0020] Example 1: Please see Figure 1 This invention provides a technical solution: a method for processing arthritis data based on ceRNA-network analysis, comprising the following steps: S1: Obtain transcriptional profile data of mRNA and non-coding RNA, trace RNA target annotation information and sample grouping labels from arthritis samples, extract candidate regulatory pairs with shared trace RNA response elements, simultaneously identify transcriptional consistency patterns between samples, jointly compare transcriptional association and target co-annotation relationship, perform regulatory association labeling, and generate a joint screening and exclusion set of transcriptional targets for arthritis. S2: Identify candidate regulatory pairs that were not marked as excluded from the joint screening of transcriptional-targeted regulation in arthritis, extract RNA molecular function annotation fields and pathway enrichment fields, match molecular function tags and signaling pathway tags, determine whether the regulatory pairs meet the conditions for enrichment of arthritis-related pathways, and generate a set of pathological pathway enrichment and functional matching verification markers. S3: Obtain candidate regulatory pairs corresponding to the pathological pathway enrichment and functional matching verification identifier set, unify the transcription unit, calculate the Pearson correlation coefficient, classify the association strength interval and regulatory confidence level, and obtain the regulatory priority determination result. S4: Based on the regulatory priority determination results, identify the set of connection edges of candidate regulatory pairs in the ceRNA network, extract the number of directly adjacent nodes and the number of indirect path nodes, integrate multi-level neighborhood structure information, inject regulatory confidence level parameters, and combine with regulatory priority to obtain the ceRNA network node regulatory impact assessment dataset. S5: Based on the ceRNA network node regulation impact assessment dataset, the network role type corresponding to the target regulatory pair molecule is matched by the node's unique identifier. The role field is parsed and loaded into the regulatory context. The ceRNA network node regulation impact assessment dataset and role type parameters are linked together. The network centrality calculation method and output format are adjusted to generate an arthritis-specific core hub node identification map.

[0021] The arthritis transcriptional targeted joint screening exclusion set includes regulatory failure status, exclusion triggering basis, and exclusion range; the pathological pathway enrichment and functional matching verification mark set includes pathway enrichment category, regulatory direction mark, and activation confidence number; the regulatory priority determination result includes association strength type, functional consistency level, and regulatory intervention level; the ceRNA network node regulatory influence assessment dataset includes direct connection weight, indirect path weight, and regulatory confidence adjustment weight; and the arthritis-specific core hub node identification map includes role type mark, number of direct connections, number of indirect paths, and regulatory influence value.

[0022] Please see Figure 2 The specific steps for generating the exclusion set for transcriptional targeting combined with screening of regulatory mechanisms in arthritis are as follows: S111: Obtain transcriptional profile data of mRNA and non-coding RNA, trace RNA targeted annotation information and sample grouping labels from arthritis samples. Extract the annotation terms belonging to the current regulatory pair according to the mapping relationship between regulatory pairs and trace RNA response elements, and establish a regulatory annotation set. The data interface program automatically established a transmission link with the gene transcription database of the China Center for Biotechnology Information. Using "Rheumatoid_Arthritis" and "Osteoarthritis" as target search terms, it located and downloaded transcriptome sequencing source files numbered GSE55235 and GSE55457. The GPL96 platform annotation library was used to map probe sequence identifiers to standard gene symbols. Matrix segmentation was performed on the original matrix to separate the messenger RNA transcription matrix from the long non-coding RNA transcription matrix. Simultaneously, a multi-threaded crawler was launched to access TargetScan and m... The iRDB database is used to batch download microRNA seed sequences and binding site information of target gene untranslated regions. A hash mapping table is constructed to store targeting relationships. All permutations of messenger RNA and non-coding RNA are traversed. For each pair, an intersection operation is performed to count the number of microRNA binding sites shared by the two. Combinations with an intersection of three or more elements are identified as initial candidates. The corresponding microRNA ID list is extracted as response element metadata. This mapping relationship and metadata are written into a dedicated table in the structured query language database to establish a regulatory annotation set, ensuring that the data source has multi-dimensional biological targeting evidence.

[0023] S112: Call the regulatory annotation set, perform cross-sample consistency test on the transcriptional profile data of mRNA and non-coding RNA, extract transcriptional trend symbols and target co-annotation status, set co-transcription threshold, perform logical AND operation, and generate regulatory feasibility judgment results; The algorithm reads the index of each record in the regulatory annotation set, locates the transcribed numerical sequences of messenger RNA and non-coding RNA in the GSE55235 transcription matrix based on the index key, calls the Spearman rank correlation operation module to perform rank transformation on the two transcribed sequences and calculates the rank correlation coefficient. Records with a correlation coefficient greater than zero and a two-sided test p-value less than 0.01 are marked as having consistent transcriptional trends. Then, the algorithm reads the microRNA support list generated in S111, counts the number of elements in the list, sets the co-transcription intensity threshold to 0.4, sets the target support quantity threshold to 3, and constructs a Boolean logic decision gate. Only when a record meets the three conditions of a correlation coefficient greater than 0.4, a p-value less than 0.01, and a co-support quantity greater than or equal to three, is the logical output of the record determined to be true. Records with a true logical result are marked as having a feasible regulatory state, and records that do not meet any of the conditions are marked as having an infeasible regulatory state. This generates a regulatory feasibility determination result containing a binary status code, completing the dual verification from statistical association to biological mechanism.

[0024] S113: Extract unique identifiers for candidate regulatory pairs, identify whether regulatory pairs lack common trace RNA target support based on the regulatory feasibility judgment results, set regulatory failure status markers for regulatory pairs lacking support, and obtain the joint screening regulatory exclusion set for transcriptional targets in arthritis. The dataset of regulatory pairs after logical judgment processing is traversed, and the universally unique identifier of each record is extracted as the primary key. For each record, the non-emptiness check and length verification of the microRNA support list are performed. Records with fewer than three elements in the support list or an empty list are identified and defined as noisy data lacking competitive endogenous RNA mechanism support. The status identifier field of such records is modified to invalid. At the same time, the unique identifier and associated metadata information of invalid records are extracted and physically migrated from the main data table to a separate exclusion table to obtain the arthritis transcriptional target joint screening regulation exclusion set. This operation achieves physical isolation of false positive regulatory relationships, ensuring that subsequent analysis is only performed on candidates with sufficient biological target evidence. The cleaned exclusion set is output as the basis for reverse filtering.

[0025] Please see Figure 3 The specific steps for generating the pathological pathway enrichment and functional matching validation label set are as follows: S211: Identify the recording information in the exclusion set of transcriptional targeting joint screening of arthritis, screen candidate regulatory pairs that are not marked as excluded, and extract the set of activatable regulatory pairs according to the mapping relationship between the regulatory pairs and arthritis-related pathways. All unique identifiers in the joint screening exclusion table for transcriptional targeting in arthritis are read, a blacklist hash set is constructed, the original regulatory annotation set is loaded, and a difference operation is performed on the original dataset using the blacklist set to retain candidate regulatory pairs in the activated state that do not appear in the blacklist. Then, the application programming interface of Kyoto Gene and Genome Encyclopedia is called to retrieve the biological pathway numbers involved by the messenger RNA in batches using the Entrez_ID as the query parameter. The retrieved pathway numbers are matched with a pre-set list of arthritis-specific pathways (including the hsa05323 rheumatoid arthritis pathway and the hsa04064 NF-kappaB signaling pathway) to screen out regulatory pairs that contain at least one arthritis-specific pathway number. Their object entities and associated attributes are stored in a set to obtain a set of activatable regulatory pairs, thus completing the targeted screening based on pathological background knowledge.

[0026] S212: Obtain the functional annotation fields and pathway enrichment fields of the corresponding mRNA and non-coding RNA of the set of activatable regulatory pairs, and generate a pathway association parameter set based on the path association between the RNA molecular ontology hierarchy and KEGG pathway number; For the set of activatable regulatory pairs, a bioinformatics annotation data extraction program was launched. Using a Perl script, the org.Hs.eg.db library in the Bioconductor package was invoked to perform batch searches for the Entrez_ID of each messenger RNA and the EnsemblID of each non-coding RNA in the set. The program traversed the molecular functional branches in the gene ontology database, extracting descriptive fields such as GO:0003723 representing RNA binding function and GO:0005515 representing protein binding function. Simultaneously, the hierarchical structure file of the Kyoto Encyclopedia of Genes and Genomes was parsed to locate the specific coordinates of each gene in the metabolic pathway map, obtaining data such as h... We constructed a multidimensional feature vector containing gene identifiers, functional terminology descriptions, and pathway attribution numbers for the tumor necrosis factor signaling pathway represented by sa04668 and the nuclear factor kappa_B signaling pathway represented by hsa04064. All retrieved unstructured annotation texts were converted into a standardized key-value pair dictionary structure. A unique association index key was established for each regulatory pair. The functional annotation fields and pathway enrichment fields were merged to generate a pathway association parameter set containing complete biological background information. This parameter set provides fundamental data support for subsequent statistical tests, ensuring the accuracy and comprehensiveness of gene function localization and eliminating the annotation bias risk caused by a single data source.

[0027] S213: Call the pathway association parameter set, extract the molecular function category in the function annotation field and the arthritis-specific pathway identifier in the pathway enrichment field, combine the hypergeometric distribution test algorithm to calculate the pathway enrichment, determine whether the regulatory pair meets the enrichment condition that the P value is less than the preset threshold, and generate a pathological pathway enrichment and function matching verification identifier set. The system retrieves data records from the pathway association parameter set, sets the background parameters required for statistical testing, defines the total number of annotated genes in the human genome as the background set total number N (valued at 20000), defines the number of genes contained in a specific arthritis pathological pathway as the target set total number M (valued at 500), and defines the total number of differentially transcribed genes involved in the currently screened activatable regulatory pairs as the sample set size n (valued at 100). It iterates through the sample set and the target set, counting the number of overlapping genes k in their intersection (valued at 15). Based on the probability density function logic of the hypergeometric distribution, it calculates the probability density function of the data in the random... The probability that exactly 15 or more genes fall into the pathological pathway when 100 genes are extracted is calculated using a precise factorial operation, yielding a P-value of 0.0032. A significance threshold of 0.05 is set, and the calculated P-value is compared with the threshold. Since 0.0032 is much smaller than 0.05, the pathway containing the regulatory pair is determined to be significantly enriched, verifying the close association between the regulatory pair and the pathological mechanism of arthritis. Based on this, a pathological pathway enrichment and functional matching verification marker set is generated, containing enrichment status indicators and P-value records, thus completing the biological statistical verification of candidate regulatory relationships.

[0028] Please see Figure 4 The specific steps for obtaining the priority determination result are as follows: S311: Obtain candidate regulatory pairs corresponding to the pathological pathway enrichment and functional matching verification identifier set, match the original transcriptome records based on the unique identifiers in the pathological pathway enrichment and functional matching verification identifier set, screen the corresponding regulatory entries, extract the regulatory pair identifier field, locate the matching record entries, screen the target regulatory pair numbers, and generate the target regulatory pair index set. The pathological pathway enrichment and functional matching validation identifier set was analyzed, and unique identifiers of validated candidate regulatory pairs were extracted. These identifiers were then parsed into corresponding probe group IDs. The original transcriptome sequencing transcription matrix file was loaded, and a hash indexing algorithm was used to quickly locate the physical storage addresses of messenger RNA probe 209983_s_at and non-coding RNA probe 231805_at within the matrix row labels. Simultaneously, the sample metadata configuration file was read, and the column index of diseased samples marked as Osteoarthritis was identified. Normal control group data was removed, retaining only the transcription values ​​of the diseased group samples. Through cross-mapping of row and column indices, local data blocks containing five biologically replicated samples were extracted from the large-scale matrix. A structured data object containing sample number, messenger RNA transcription value, and non-coding RNA transcription value was constructed. This data object was encapsulated as a target regulatory pair index set, providing cleaned and aligned pure numerical input for subsequent correlation calculations, ensuring that the numerical calculation process is performed only on sample data under the target pathological state.

[0029] S312: Call the target regulatory pair index set, extract the transcriptional vectors of mRNA and non-coding RNA for each corresponding target regulatory pair, and simultaneously identify the sample grouping field using the following formula: Calculate the Pearson correlation coefficient between two transcription vectors and generate a set of correlation coefficients; in, Indicates mRNA in the first Transcription amount in each sample Indicates that non-coding RNA is in the first Transcription amount in each sample; and These represent the average transcription levels of mRNA and non-coding RNA across all samples, respectively. For the sample size, This represents the Pearson correlation coefficient, and its value range is... ,when This indicates that the two are positively correlated. Indicates a negative correlation. This indicates that there is no linear correlation between the two; Based on the target regulatory pair index set, specific transcription values ​​were retrieved from the database. The target regulatory pair numbered Pair_101 was selected for analysis, and its mRNA transcription vectors in five arthritis samples were extracted. and the corresponding non-coding RNA transcription vector As shown in Table 1, the average amount of mRNA transcribed was first calculated. Average transcription amount of non-coding RNA Next, the sum of the products of biases for each sample is calculated. and the product of the sum of squares of their respective deviations. Finally, the Pearson correlation coefficient was calculated using the formula: The advantage of the formula is that it eliminates the influence of differences in transcription magnitude between different RNA molecules through standardized covariance calculation, and accurately quantifies the tightness of linear co-transcription between the two. In this example, the first step is to calculate the numerator: Sample 1: ; Sample 2: ; Sample 3: ; Sample 4: ; Sample 5: ; Total of numerators ; The second step is to calculate the denominator: Sum of squared deviations ; Sum of squared deviations ; denominator ; The third step is to calculate the G value: ; The results indicate that mRNA and non-coding RNA in Pair_101 show a strong positive correlation in arthritis samples, consistent with the co-transcriptional characteristics caused by sponge adsorption in the ceRNA mechanism, and the generated correlation coefficient set includes this value. Table 1: Data on mRNA and non-coding RNA transcription levels in arthritis samples Table 1 lists the specific transcription values ​​for the five samples used to calculate the Pearson correlation coefficient.

[0030] S313: Compare each coefficient in the set of correlation coefficients with the preset correlation strength range to determine whether the correlation strength of the target control pair is within the preset range, and generate the control priority determination result. The program iterates through each calculation result in the set of correlation coefficients, loads a preset correlation strength grading standard, sets the lower limit of the strong correlation interval to 0.7 and the upper limit to 1.5, the lower limit of the medium correlation interval to 0.4 and the upper limit to 0.7, and the weak correlation interval to less than 0.4. For the correlation coefficient of 0.998 calculated above, the program compares it with the boundary values ​​of each interval, determines whether the value falls within the strong correlation interval, and marks the corresponding target control pair as a first-level priority. If the calculation result is 0.55, it falls into the medium correlation interval and is marked as a second-level priority. If the result is 0.2, it is marked as a third-level priority. Through multi-level judgment logic, the continuous correlation coefficient values ​​are converted into discrete priority classification labels, generating a control priority judgment result containing the classification results. This result directly determines the edge weight allocation strategy in the subsequent network construction process, ensuring that high-confidence control relationships dominate in network analysis.

[0031] Please see Figure 5 The specific steps for obtaining the dataset for assessing the regulatory impact of ceRNA network nodes are as follows: S411: Based on the regulatory priority determination results, analyze the adjacency relationship of the target regulatory pair nodes in the ceRNA network, locate their position in the global network according to the node's unique identifier, and extract the list of directly adjacent nodes and the list of two-hop adjacent nodes to generate a neighborhood structure dataset. Based on the priority determination results, a ceRNA regulatory network is constructed using an initialized graph database. Starting from the mRNA node, a breadth-first search algorithm is executed to traverse the network edges and identify all nodes directly connected to the starting point, defining them as one-hop neighbors. For example, node M1 connects to nodes L1 and L2. Then, starting from the one-hop neighbors, the search continues outward to identify nodes connected to one-hop neighbors but not visited, defining them as two-hop neighbors. For example, node L1 connects to node M2. The connections between M1, L1, L2, M2, and each other are extracted, and a local subgraph structure containing node ID, connection level, and connection type is constructed. This subgraph structure is serialized and stored to generate a neighborhood structure data set. This data set completely preserves the topological environment information of the target node in the local network, providing a topological foundation for subsequent influence assessment based on neighborhood features.

[0032] S412: Obtain the neighborhood structure data set, identify and regulate the association confidence level, and perform weighted labeling on the number of direct adjacencies and the number of indirect paths based on it to generate a weighted neighborhood parameter set. The neighborhood structure data set and its corresponding priority labels are read, and the connection edges in the network are weighted. For connection edges marked as first-level priority, a weight value of 1.0 is assigned, and for connection edges marked as second-level priority, a weight value of 0.5 is assigned. The neighborhood list of the target node is traversed, and the number of directly connected one-hop neighbors is counted as the direct degree index, and the number of two-hop neighbors is counted as the indirect path index. The weight values ​​and the indexes are associated and bound. For example, the connection between node T and neighbor N1 is a first-level connection, and the weight is recorded as 1.0. The connection between node T and neighbor N2 is a second-level connection, and the weight is recorded as 0.5. At the same time, the degree centrality value of each neighbor node is recorded. All the quantified parameters are encapsulated to generate a weighted neighborhood parameter set, as shown in Table 2. This parameter set quantifies the strength of each connection in the local network and the topological importance of each node, providing a weighted numerical input for accurately calculating the influence of node regulation.

[0033] Table 2: Example Table of Weighted Neighborhood Parameter Groups Table 2 lists the specific values ​​of various neighborhood parameters required for impact assessment of target nodes.

[0034] S413: Call the weighted neighborhood parameter set, perform a linear combination operation based on the combined distribution characteristics of direct adjacency weights and indirect path weights in the current weighted state, using the formula: : Calculate the regulatory influence value to obtain a dataset assessing the regulatory influence of ceRNA network nodes; in, Represents the influence value of regulation. Representing the The adjacency weight between each node and the target node; Representing the The weight of the direct adjacency path from each node to the target node; Representing the Adjustment factor for each node, Represents a regulatory factor. This represents the total number of associated nodes participating in the calculation within the ceRNA network; The weighted neighborhood parameter set is invoked to calculate the core metric, and the target node Node_T is selected for evaluation. It is assumed that this node has two associated nodes in the network participating in the calculation. ), which are Node_1 and Node_2 respectively. The adjacency weights of Node_1 and the target node are obtained from the parameter set. (Representing strong correlation), weight of directly adjacent paths (Representing a direct connection), the degree of Node_1 itself is the adjustment factor. Get the adjacency weight of Node_2. (Representing moderate relevance), path weight (Representing a two-hop connection), the adjustment factor for Node_2 ; Set adjustment factor To control the nonlinear contribution of node degree to influence; Calculate the influence value of regulation using the formula: ; The advantage of this formula lies in its comprehensive consideration of the connectivity strength of neighboring nodes, distance decay, and the topological importance of the neighbors themselves, through adjustment factors. Balance the influence weights of high-degree and low-degree nodes; The specific calculations are as follows: The first step is to calculate the contribution of Node_1: Term_1 ; The second step is to calculate the contribution of Node_2: Term_2 ; The third step is to perform linear combination summation: ; The results indicate that Node_T has a comprehensive regulatory influence of 2.75 in the current local network environment. This value will be stored in the dataset to obtain the ceRNA network node regulatory influence assessment dataset, which will serve as a quantitative basis for subsequent judgment on whether the node is a core hub.

[0035] Please see Figure 6 The specific steps for obtaining the arthritis-specific core hub node identification map are as follows: S511: Based on the ceRNA network node regulation impact assessment dataset, identify the correspondence between the node's unique identifier and the target regulatory pair molecule, filter the role type records pointed to by the current regulatory network, and generate a target role type set; The program loads all influence values ​​from the ceRNA network node regulation impact assessment dataset, executes a quicksort algorithm to sort the values ​​in descending order, calculates the percentile statistics of the entire value distribution, sets the criteria for determining a core hub role as being in the top 10% of the distribution, the criteria for a bottleneck bridging role as being in the 10% to 30% range, and the criteria for a marginal moderating role as being in the bottom 70% range. For the calculated influence value of 3.457, if the value is in the top 5% of the overall distribution, the program will identify the corresponding node as a core hub role according to the criteria; if the value is in the top 15%, it will be identified as a bottleneck bridging role. Through this process, continuous influence values ​​are mapped to discrete role function labels, generating a target role type set and clarifying the functional positioning of each node in the network.

[0036] S512: Based on the target role type set, and according to the alignment relationship between role type identifier and regulatory priority, determine the centrality computational requirements of the corresponding role type based on the ceRNA network node regulatory impact assessment dataset, and generate a centrality parameter allocation group; Based on the classification results in the target role type set, differentiated centrality calculation parameters are configured for nodes of different roles. For nodes identified as core hub roles, a full centrality calculation strategy is configured, including degree centrality, betweenness centrality, proximity centrality, and eigenvector centrality algorithms. The tolerance parameter for computational convergence is set to 10 to the power of -6 to ensure high accuracy. For nodes of edge adjustment roles, only a basic degree centrality calculation strategy is configured, and the tolerance parameter is set to 10 to the power of -4 to reduce computational resource consumption. The algorithm list and parameter configuration corresponding to each role are encapsulated into an instruction package to generate a centrality parameter allocation group. This allocation group realizes the on-demand allocation of computational resources, which optimizes the overall computational efficiency while ensuring the accuracy of core node evaluation.

[0037] S513: Write the calibrated calculation parameters in the centrality parameter allocation group into the centrality algorithm of the corresponding role type, identify the output state after the centrality calculation adjustment, and output the arthritis-specific core hub node identification map; The centrality parameter allocation group is imported into the graph theory analysis engine. High-precision multi-dimensional centrality calculations are performed on the nodes marked as core hubs to obtain their accurate topological index values. Combined with the strong correlation results obtained from S312 and the pathological pathway enrichment status verified by S213, multi-condition joint screening is performed to retain only nodes that simultaneously satisfy high centrality, high correlation and enrichment in arthritis-specific pathways. The visualization rendering module is called to map the selected nodes as key entities in the atlas. The display size of the nodes is set to be proportional to their centrality value, and the color of the nodes is set to bright red for distinction. The background network nodes are set to gray. Finally, an arthritis-specific core hub node identification atlas that intuitively displays key pathogenic molecules is generated. This atlas directly reveals the core regulatory elements in the disease network and provides precise biological targets for subsequent drug target development.

[0038] The ceRNA-network analysis-based arthritis data processing system is used to execute the above-mentioned ceRNA-network analysis-based arthritis data processing method. The system includes: The regulation exclusion identification module obtains the transcriptional profile data of mRNA and non-coding RNA, trace RNA targeting annotation information and sample grouping labels in arthritis samples, and jointly judges whether the current regulatory pair meets the ceRNA action conditions, and generates an arthritis transcriptional targeting joint screening regulation exclusion set. The regulation activation judgment module identifies candidate regulatory pairs that were not marked as excluded from the joint screening of transcriptional targets in arthritis. It extracts RNA molecular functional annotation fields and pathway enrichment fields, determines whether the conditions for arthritis-related pathway enrichment are met, and generates a set of pathological pathway enrichment and functional matching verification markers. The regulatory priority analysis module calculates the correlation between the transcription vectors of mRNA and non-coding RNA based on the candidate regulatory pairs corresponding to the pathological pathway enrichment and functional matching verification mark set, determines the correlation strength of the regulatory pairs, and obtains the regulatory priority determination results. The regulatory influence calculation module identifies the adjacency relationship of candidate regulatory pairs in the ceRNA network based on the regulatory priority determination results, and calculates the ceRNA network node regulatory influence assessment dataset by combining the regulatory association confidence level as a regulation term. The key node identification module matches the network role type corresponding to the target regulatory pair molecules based on the node's unique identifier. It adjusts the execution parameters of the centrality algorithm according to the ceRNA network node regulation influence assessment dataset to generate an arthritis-specific core hub node identification map.

[0039] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A method for processing arthritis data based on ceRNA-network analysis, characterized in that, Includes the following steps: S1: Obtain transcriptional profile data of mRNA and non-coding RNA, trace RNA target annotation information and sample grouping labels from arthritis samples, extract candidate regulatory pairs with shared trace RNA response elements, simultaneously identify transcriptional consistency patterns between samples, jointly compare transcriptional association and target co-annotation relationship, perform regulatory association labeling, and generate a joint screening and exclusion set of transcriptional targets for arthritis. S2: Based on the arthritis transcription-targeted joint screening regulatory exclusion set, extract RNA molecular function annotation fields and pathway enrichment fields, match molecular function tags and signaling pathway tags, and generate a pathological pathway enrichment and functional matching verification identifier set. S3: Obtain candidate regulatory pairs corresponding to the pathological pathway enrichment and functional matching verification identifier set, unify the transcription unit, classify the association strength interval and regulatory confidence level, and obtain the regulatory priority determination result. S4: Based on the regulatory priority determination results, identify the set of connection edges of candidate regulatory pairs in the ceRNA network, extract the number of directly adjacent nodes and the number of indirect path nodes, integrate multi-level neighborhood structure information, inject regulatory confidence level parameters, and obtain the ceRNA network node regulatory impact assessment dataset.

2. The arthritis data processing method based on ceRNA-network analysis according to claim 1, characterized in that, The arthritis transcription-targeted joint screening and regulatory exclusion set includes regulatory failure status, exclusion triggering basis, and exclusion range. The pathological pathway enrichment and functional matching verification identifier set includes pathway enrichment category, regulatory direction identifier, and activation confidence number. The regulatory priority determination result includes association strength type, functional consistency level, and regulatory intervention level. The ceRNA network node regulatory impact assessment dataset includes direct connection weight, indirect path weight, and regulatory confidence adjustment weight.

3. The method for processing arthritis data based on ceRNA-network analysis according to claim 1, characterized in that, The specific steps for generating the exclusion set for the joint screening of transcriptional targeting in arthritis are as follows: S111: Obtain transcriptional profile data of mRNA and non-coding RNA, trace RNA targeted annotation information and sample grouping labels from arthritis samples. Extract the annotation terms belonging to the current regulatory pair according to the mapping relationship between regulatory pairs and trace RNA response elements, and establish a regulatory annotation set. S112: Call the regulatory annotation set, perform cross-sample consistency test on the transcriptional profile data of mRNA and non-coding RNA, extract transcriptional trend symbols and target co-annotation status, set co-transcription threshold, perform logical AND operation, and generate regulatory feasibility judgment result; S113: Extract unique identifiers for candidate regulatory pairs, identify whether the regulatory pairs lack common trace RNA targeting support based on the regulatory feasibility determination results, set regulatory failure status markers for regulatory pairs lacking support, and obtain the joint screening regulatory exclusion set for arthritis transcriptional targeting.

4. The arthritis data processing method based on ceRNA-network analysis according to claim 3, characterized in that, The specific steps for generating the pathological pathway enrichment and functional matching verification identifier set are as follows: S211: Identify the recorded information in the exclusion set of the transcriptional targeting joint screening regulation of arthritis, screen candidate regulatory pairs that are not marked as excluded, and extract the set of activatable regulatory pairs according to the mapping relationship between the regulatory pairs and the arthritis-related pathways. S212: Obtain the functional annotation fields and pathway enrichment fields of the corresponding mRNA and non-coding RNA of the set of activatable regulatory pairs, and generate a pathway association parameter set based on the path association between the RNA molecular ontology hierarchy and KEGG pathway number; S213: Call the pathway association parameter set, extract the molecular function category in the function annotation field and the arthritis-specific pathway identifier in the pathway enrichment field, calculate the pathway enrichment by combining the hypergeometric distribution test algorithm, determine whether the regulatory pair meets the enrichment condition that the P value is less than the preset threshold, and generate a pathological pathway enrichment and function matching verification identifier set.

5. The arthritis data processing method based on ceRNA-network analysis according to claim 4, characterized in that, The specific steps for obtaining the control priority determination result are as follows: S311: Obtain candidate regulatory pairs corresponding to the pathological pathway enrichment and functional matching verification identifier set, match the original transcriptome records according to the unique identifiers in the pathological pathway enrichment and functional matching verification identifier set, screen the corresponding regulatory entries, extract the regulatory pair identifier field, locate the matching record entries, screen the target regulatory pair numbers, and generate the target regulatory pair index set. S312: Call the target regulatory pair index set, extract the transcription vectors of mRNA and non-coding RNA in each corresponding target regulatory pair, and simultaneously identify the sample grouping field, calculate the Pearson correlation coefficient between the two transcription vectors, and generate a set of correlation coefficients. S313: Compare each coefficient in the set of correlation coefficients with the preset correlation strength range to determine whether the correlation strength of the target control pair is within the preset range, and generate a control priority determination result.

6. The method for processing arthritis data based on ceRNA-network analysis according to claim 5, characterized in that, The specific steps for obtaining the dataset for assessing the regulatory impact of ceRNA network nodes are as follows: S411: Based on the regulatory priority determination result, analyze the adjacency relationship of the target regulatory pair node in the ceRNA network, locate its position in the global network according to the node's unique identifier, and extract the list of directly adjacent nodes and the list of two-hop adjacent nodes to generate a neighborhood structure data set. S412: Obtain the neighborhood structure data group, identify and regulate the association confidence level, and perform weighted labeling processing on the number of direct adjacencies and the number of indirect paths based on it to generate a weighted neighborhood parameter group. S413: Call the weighted neighborhood parameter group, and perform a linear combination operation based on the combined distribution characteristics of direct adjacency weights and indirect path weights in the current weighted state, using the formula: : Calculate the regulatory influence value to obtain a dataset assessing the regulatory influence of ceRNA network nodes; in, Represents the value of regulatory influence. Representing the The adjacency weight between each node and the target node; Representing the The weight of the direct adjacency path from each node to the target node; Representing the Adjustment factor for each node, Represents a regulatory factor. This represents the total number of associated nodes participating in the calculation within the ceRNA network.

7. The method for processing arthritis data based on ceRNA-network analysis according to claim 1, characterized in that, The method also includes step S5: S5: Based on the ceRNA network node regulation impact assessment dataset, the network role type corresponding to the target regulatory pair molecule is matched by the node's unique identifier. The role field is parsed and loaded into the regulatory context. The ceRNA network node regulation impact assessment dataset and role type parameters are linked together. The network centrality calculation method and output format are adjusted to generate an arthritis-specific core hub node identification map. The arthritis-specific core hub node identification map includes role type identifiers, number of direct connections, number of indirect paths, and regulatory influence value.

8. The method for processing arthritis data based on ceRNA-network analysis according to claim 7, characterized in that, The specific steps for obtaining the arthritis-specific core hub node identification map are as follows: S511: Based on the ceRNA network node regulatory influence assessment dataset, identify the correspondence between the node's unique identifier and the target regulatory pair molecule, filter the role type records pointed to by the current regulatory network, and generate a target role type set; S512: Based on the target role type set, and according to the alignment relationship between role type identifier and regulatory priority, determine the centrality computational requirement of the corresponding role type based on the ceRNA network node regulatory impact assessment dataset, and generate a centrality parameter allocation group; S513: Write the calibrated calculation parameters in the centrality parameter allocation group into the centrality algorithm of the corresponding role type, identify the output state after the centrality calculation adjustment, and output the arthritis-specific core hub node identification map.

9. A data processing system for arthritis based on ceRNA-network analysis, characterized in that, The system is used to implement the arthritis data processing method based on ceRNA-network analysis according to any one of claims 1-8, and the system comprises: The regulation exclusion identification module obtains the transcriptional profile data of mRNA and non-coding RNA, trace RNA targeting annotation information and sample grouping labels in arthritis samples, and jointly judges whether the current regulatory pair meets the ceRNA action conditions, and generates an arthritis transcriptional targeting joint screening regulation exclusion set. The regulation activation judgment module identifies candidate regulatory pairs that are not marked as excluded from the joint screening regulation exclusion set of the arthritis transcriptional target, extracts RNA molecular functional annotation fields and pathway enrichment fields, determines whether the conditions for arthritis-related pathway enrichment are met, and generates a pathological pathway enrichment and functional matching verification identifier set. The regulatory priority analysis module calculates the correlation between the transcriptional vectors of mRNA and non-coding RNA based on the candidate regulatory pairs corresponding to the pathological pathway enrichment and functional matching verification identifier set, determines the correlation strength of the regulatory pairs, and obtains the regulatory priority determination result. Based on the regulatory priority determination results, the regulatory influence calculation module identifies the adjacency relationship of candidate regulatory pairs in the ceRNA network, and combines the regulatory association confidence level as a regulation term to calculate the ceRNA network node regulatory influence assessment dataset. The key node identification module matches the network role type corresponding to the target regulatory pair molecules based on the node's unique identifier. It adjusts the execution parameters of the centrality algorithm according to the ceRNA network node regulation influence assessment dataset to generate an arthritis-specific core hub node identification map.