Macromolecule analysis data sharing management method and system

By cleaning and standardizing macromolecular analysis data, and combining domain identification and dynamic key technology, the problems of biological relevance and access control in macromolecular data sharing are solved, achieving high-quality data sharing with both security and efficiency.

CN120636557BActive Publication Date: 2025-11-18SHANGHAI TAICHU BIOTECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511114174.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-18
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

Current methods for sharing and managing macromolecular analysis data have significant shortcomings in terms of the biological relevance of data preprocessing, the precision of access control, and the completeness of validity verification, which fail to meet the needs of life science research for high-quality data sharing.

Method used

By cleaning and standardizing the original sequence data of macromolecules, combining domain identification tools to divide the domain segments and flexible region segments, calculating the functional sensitivity index and user domain matching degree, generating dynamic keys for access control, and ensuring the security and accuracy of data transmission through data integrity and homology verification.

Benefits of technology

It enables precise access control for macromolecular analysis data, ensuring that users can access data highly relevant to their research directions, reducing the risk of leakage of core biological information, improving the security and efficiency of data sharing, and providing reliable data support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636557B_ABST
    Figure CN120636557B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of optical character recognition, and discloses a macromolecule analysis data sharing management method and system, which comprises the following steps: converting an original sequence into a standardized sequence based on residue weight and a standardization formula; dividing the standardized sequence into a plurality of domain fragments and a plurality of flexible region fragments, and calculating the functional sensitivity index of the current domain fragment and the user field matching degree; generating a dynamic key; obtaining a recombined sequence; verifying the recombined sequence, and if the second preset condition is met, determining that the recombined sequence is valid, and sending the recombined sequence to the client of the current user. Through implementation of the application, the problem that the current macromolecule analysis data sharing management method has significant deficiencies in biological correlation in data preprocessing, precision in authority control, integrity in effectiveness verification and the like, and cannot meet the demand of life science research on high-quality data sharing is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of optical character recognition technology, specifically to a method and system for sharing and managing macromolecular analysis data. Background Technology

[0002] With the rapid development of molecular biology and genomics, macromolecular sequence data is growing exponentially. This data, containing key biological characteristics such as genetic information and functional mechanisms, has become a core resource in life science research, drug development, and disease diagnosis. However, the sharing of macromolecular analysis data faces a prominent contradiction between security and availability: on the one hand, the high-value biological information contained in the data needs strict protection to prevent unauthorized access or misuse; on the other hand, the deepening of scientific collaboration requires data to circulate efficiently within compliant scope, providing support for cross-institutional and cross-disciplinary research. Currently, macromolecular data sharing mainly adopts three models: First, full open sharing, which involves making complete sequence data publicly available to all users through public databases. While this model ensures data availability, it cannot achieve fine-grained access control, leading to the risk of sensitive information leakage. Second, role-based access control, which assigns data access permissions by pre-defining user roles. However, this static permission mechanism does not consider the biological characteristics of the data itself, potentially leading to users accessing highly sensitive fragments unrelated to their research, or being unable to access necessary functional region data due to excessive access restrictions, thus reducing the accuracy of sharing. Thirdly, there is the encrypted transmission mode, which uses symmetric or asymmetric encryption technology to protect the data transmission process. However, the integrity and biological validity of the decrypted data lack a verification mechanism. If fragments are lost or tampered with during transmission, subsequent analysis may lead to incorrect conclusions.

[0003] Current methods for sharing and managing macromolecular analysis data have significant shortcomings in terms of the biological relevance of data preprocessing, the precision of access control, and the completeness of validity verification, which fail to meet the needs of life science research for high-quality data sharing. Summary of the Invention

[0004] In view of this, the present invention provides a method and system for sharing and managing macromolecular analysis data, in order to solve the problem that current macromolecular analysis data sharing and management methods have significant shortcomings in terms of the biological relevance of data preprocessing, the precision of access control, and the completeness of validity verification, and cannot meet the needs of life science research for high-quality data sharing.

[0005] In a first aspect, the present invention provides a method for sharing and managing macromolecular analysis data. This method includes: acquiring and cleaning the raw sequence data of the macromolecule, and performing a standardization transformation to obtain a standardized sequence; using a domain identification tool to determine the domain boundaries of the macromolecule, and dividing the standardized sequence into several domain segments and several flexible region segments based on the domain boundaries and secondary structure prediction results; acquiring the functional sensitivity index, conservation score, functional importance score, and pathway association score of each domain segment, and calculating the functional sensitivity index of the current domain segment; acquiring the current user's research domain keyword set, the functional tags of each domain segment, and semantic similarity... The total number of functional tags is calculated, and the user domain matching degree of the current structural domain segment is calculated. If the functional sensitivity index and user domain matching degree of the current structural domain segment meet the first preset condition, the current user is granted access to the current structural domain segment. Based on the authorized current structural domain segment, the conservatism score, the timestamp, and the user's random number, a dynamic key is generated. Each structural domain segment is concatenated according to its position order in the original sequence data to obtain a reconstructed sequence. The data integrity of the reconstructed sequence is verified, and the homology score between the reconstructed sequence and the original sequence is calculated. If the data integrity verification result and the homology score meet the second preset condition, the reconstructed sequence is determined to be valid, and the reconstructed sequence is sent to the current user's client.

[0006] The macromolecule analysis data sharing and management method provided in this embodiment firstly cleans the original macromolecule sequence data and calculates residue weights by combining the base weights of residues and conservation scores, then converts it into a standardized sequence. This process not only removes interference information from low-quality regions in the original data, improving data reliability, but also incorporates the biological characteristics of the sequence into the standardization process through residue weights, ensuring that the standardized sequence not only has data consistency but also retains key biological information. Secondly, a domain identification tool is used to determine domain boundaries, and the results of secondary structure prediction are combined to divide the data into domain fragments and flexible region fragments. Each fragment contains sequence data and biological metadata, ensuring that the divided fragments conform to the natural structure and functional characteristics of the macromolecule. The integrity of the domain fragments is maintained, avoiding functional unit destruction caused by random segmentation, while the biological metadata provides direct biological basis for subsequent permission determination and data analysis. Then, by acquiring the functional sensitivity index, conservation score, functional importance score, and pathway association score of the domain fragments and calculating the functional sensitivity index, the biological functional sensitivity of each domain fragment is quantified, clarifying its importance in macromolecular function execution. This allows for more stringent protection of highly sensitive and important fragments, reducing the risk of leakage of core biological information. Furthermore, by combining the user's research domain keyword set with the functional tags of the domain fragments, and using semantic similarity and the total number of functional tags to calculate the user domain matching degree, the relevance between user needs and fragment functions can be accurately measured. This ensures that the fragments obtained by users are highly relevant to their research direction, reducing the transmission and access of irrelevant data and improving the targeting and efficiency of data sharing. Finally, by combining the functional sensitivity index and user domain matching degree as authorization conditions, a two-dimensional access control system is implemented. This considers both the inherent sensitivity of the domain fragments and the actual needs of user research, avoiding excessive openness or restriction of permissions caused by single-dimensional authorization. While ensuring data security, it also ensures that legitimate users can obtain the data they need, balancing security and usability. Furthermore, by generating dynamic keys based on authorized structural domain fragments, conservatism scores, timestamps, and user-generated random numbers, the keys are tightly bound to data characteristics, user identity, and time. This dynamic association mechanism significantly improves the uniqueness and timeliness of the keys, reduces the risk of key reuse or forgery, and provides higher security for data transmission.

[0007] Then, the recombinant sequence is obtained by splicing authorized fragments, dynamic keys, and domain boundaries. The positional information of the domain boundaries ensures the accuracy of fragment splicing, making the recombinant sequence structurally close to the arrangement of the original sequence. Simultaneously, the decryption and verification process of the dynamic key further ensures the authenticity and integrity of the fragments used for splicing. Finally, by verifying the data integrity of the domain fragments and calculating the homology score between the recombinant and original sequences, and using a dual-condition determination of the recombinant sequence's validity, the integrity verification ensures that the data has not been tampered with or lost during transmission and splicing, and the homology score verifies the biological consistency between the recombinant and original sequences. This ensures that the sequence ultimately obtained by the user is not only formally complete but also possesses genuine biological functional characteristics, providing reliable data support for subsequent macromolecular analysis research. In summary, the macromolecular analysis data sharing management method provided in this embodiment effectively solves the significant shortcomings of current macromolecular analysis data sharing management methods in terms of the biological relevance of data preprocessing, the precision of access control, and the completeness of validity verification, failing to meet the needs of life science research for high-quality data sharing.

[0008] In one optional implementation, the process of obtaining and cleaning the raw sequence data of the macromolecule and performing a standardization transformation to obtain a standardized sequence includes:

[0009] The sliding window size is set to 5 residues. When the proportion of N bases in the sequence region within the window exceeds a preset threshold, the above sequence region is marked as a low-quality region and the region coordinates are recorded.

[0010] The residue weights are calculated based on the baseline weights and conservation scores of the residues in the original sequence data, using the following formula:

[0011]

[0012] Among them, For residues Biological weights, As the base value for weights,

[0013] The conservation score of residues in homologous sequences,

[0014] The original sequence is converted into a normalized sequence based on residue weights and a normalization formula, which is as follows:

[0015]

[0016] in, For the first , For residues The position index in the sequence, where L is the length of the original sequence. For residues Biological weights, This is the quality correction factor.

[0017] In one optional implementation, the domain boundaries of the macromolecule are determined using a domain identification tool, and the normalized sequence is divided into several domain segments and several flexible region segments based on the domain boundaries and secondary structure prediction results, including:

[0018] The HMMER tool was used to compare the data against the Pfam database to identify domain boundaries in the normalized sequences. ,

[0019] ,in, Indicates the starting position of the m-th structural domain. and termination Location, for satisfying Annotate the known structural domains within the structural domains;

[0020] For unannotated regions, secondary structures are predicted using PSIPRED. When the number of consecutive identical secondary structure residues is ≥8, the corresponding unannotated region is marked as a potential domain. ;

[0021] The resulting structural domain fragments are obtained:

[0022]

[0023] in, For the standardized sequence fragments corresponding to the structural domains, This is biometric metadata, which includes domain IDs, secondary structure proportions, and functional tags for GO annotations.

[0024] Several flexible region segments are obtained by splitting based on the flexibility index, splitting coefficient, and segment length. The flexibility index... for: ,in, The hydrophobicity index of residues, the resolution coefficient for: Fragment length for: , where s and e are the starting and ending positions of the flexible region, respectively.

[0025] In one optional implementation, the functional sensitivity index, conservation score, functional importance score, and pathway correlation score of each domain segment are obtained, and the functional sensitivity index of the current domain segment is calculated using the following formula:

[0026]

[0027] in, Let be the functional sensitivity index of the m-th structural domain segment. Let m be the conservatism score of the m-th structural domain segment. The functional importance score for the m-th structural domain segment. Let α+β+γ be the pathway connectivity score for the m-th domain segment, where α+β+γ=1.

[0028] In one optional implementation, the functional sensitivity index, conservation score, functional importance score, and pathway correlation score of each domain segment are obtained, and the functional sensitivity index of the current domain segment is calculated using the following formula:

[0029]

[0030] in, For user domain matching degree, This is a collection of keywords for current user research. This refers to the t-th functional label of the m-th structural domain segment; Let m be the semantic similarity of the m-th structural domain segment, based on The ontology calculation shows that T is the total number of functional tags for the m-th structural domain segment.

[0031] In one optional implementation, the first preset condition is:

[0032]

[0033] in, This is the permission threshold.

[0034] In an alternative implementation, the dynamic key is generated based on the authorized current domain fragment, the conservatism score, the timestamp, and the user's random number, as shown in the following formula:

[0035]

[0036] in, For the current user Visit the Dynamic keys for each structural domain segment For the first A standardized sequence fragment of a structural domain segment, Let m be the conservatism score of the m-th structural domain segment. Millisecond-level timestamps For the current user A unique random number is generated during registration. For hash functions, This is for string concatenation operations.

[0037] In one optional implementation, the above data integrity verification formula is as follows:

[0038]

[0039] in, Let m be the checksum of the m-th structural domain segment. Let i be the normalized value of the i-th residue. , Let be the start and end positions of the m-th structural domain segment.

[0040] In one optional implementation, the above-mentioned homology score is calculated using the following formula:

[0041]

[0042] Where H is the homology score, with a value ranging from 0 to 1. This represents the number of matching residues between the recombinant sequence and the original sequence. The length of the original sequence. This represents the length of the recombinant sequence.

[0043] Secondly, this invention provides a macromolecule analysis data sharing and management system, which includes: a standardization module for acquiring and cleaning the raw sequence data of macromolecules, and performing standardization transformation to obtain standardized sequences; a partitioning module for determining the domain boundaries of the macromolecules using domain identification tools, and dividing the standardized sequences into several domain segments and several flexible region segments based on the domain boundaries and secondary structure prediction results; a first calculation module for acquiring the functional sensitivity index, conservation score, functional importance score, and pathway association score of each domain segment, and calculating the functional sensitivity index of the current domain segment; and a second calculation module for acquiring the current user's research field keyword set, the functional labels of each domain segment, and semantic correlation... The system comprises the following modules: a similarity module and a functional tag count module; an access module, which grants the current user access to the current structural domain segment if the functional sensitivity index and user domain matching degree of the current structural domain segment meet a first preset condition; a generation module, which generates a dynamic key based on the authorized current structural domain segment, the conservatism score, the timestamp, and the user's random number; a splicing module, which splices the structural domain segments according to their positional order in the original sequence data to obtain a recombined sequence; and a sharing module, which verifies the data integrity of the recombined sequence, calculates the homology score between the recombined sequence and the original sequence, and determines the recombined sequence to be valid if the data integrity verification result and the homology score meet a second preset condition, and sends the recombined sequence to the current user's client. Attached Figure Description

[0044] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0045] Figure 1 This is a flowchart illustrating a macromolecular analysis data sharing and management method according to an embodiment of the present invention;

[0046] Figure 2 This is a structural block diagram of a macromolecular analysis data sharing management system according to an embodiment of the present invention. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0048] With the rapid development of molecular biology and genomics, macromolecular sequence data is growing exponentially. This data, containing key biological characteristics such as genetic information and functional mechanisms, has become a core resource in life science research, drug development, and disease diagnosis. However, the sharing of macromolecular analysis data faces a prominent contradiction between security and availability: on the one hand, the high-value biological information contained in the data needs strict protection to prevent unauthorized access or misuse; on the other hand, the deepening of scientific collaboration requires data to circulate efficiently within compliant scope, providing support for cross-institutional and cross-disciplinary research. Currently, macromolecular data sharing mainly adopts three models: First, full open sharing, which involves making complete sequence data publicly available to all users through public databases. While this model ensures data availability, it cannot achieve fine-grained access control, leading to the risk of sensitive information leakage. Second, role-based access control, which assigns data access permissions by pre-defining user roles. However, this static permission mechanism does not consider the biological characteristics of the data itself, potentially leading to users accessing highly sensitive fragments unrelated to their research, or being unable to access necessary functional region data due to excessive access restrictions, thus reducing the accuracy of sharing. Thirdly, there is the encrypted transmission mode, which uses symmetric or asymmetric encryption technology to protect the data transmission process. However, the integrity and biological validity of the decrypted data lack a verification mechanism. If fragments are lost or tampered with during transmission, subsequent analysis may lead to incorrect conclusions.

[0049] Current methods for sharing and managing macromolecular analysis data have significant shortcomings in terms of the biological relevance of data preprocessing, the precision of access control, and the completeness of validity verification, which fail to meet the needs of life science research for high-quality data sharing.

[0050] The macromolecule analysis data sharing and management method provided in this embodiment firstly cleans the original macromolecule sequence data and calculates residue weights by combining the base weights of residues and conservation scores, then converts it into a standardized sequence. This process not only removes interference information from low-quality regions in the original data, improving data reliability, but also incorporates the biological characteristics of the sequence into the standardization process through residue weights, ensuring that the standardized sequence not only has data consistency but also retains key biological information. Secondly, a domain identification tool is used to determine domain boundaries, and the results of secondary structure prediction are combined to divide the data into domain fragments and flexible region fragments. Each fragment contains sequence data and biological metadata, ensuring that the divided fragments conform to the natural structure and functional characteristics of the macromolecule. The integrity of the domain fragments is maintained, avoiding functional unit destruction caused by random segmentation, while the biological metadata provides direct biological basis for subsequent permission determination and data analysis. Then, by acquiring the functional sensitivity index, conservation score, functional importance score, and pathway association score of the domain fragments and calculating the functional sensitivity index, the biological functional sensitivity of each domain fragment is quantified, clarifying its importance in macromolecular function execution. This allows for more stringent protection of highly sensitive and important fragments, reducing the risk of leakage of core biological information. Furthermore, by combining the user's research domain keyword set with the functional tags of the domain fragments, and using semantic similarity and the total number of functional tags to calculate the user domain matching degree, the relevance between user needs and fragment functions can be accurately measured. This ensures that the fragments obtained by users are highly relevant to their research direction, reducing the transmission and access of irrelevant data and improving the targeting and efficiency of data sharing. Finally, by combining the functional sensitivity index and user domain matching degree as authorization conditions, a two-dimensional access control system is implemented. This considers both the inherent sensitivity of the domain fragments and the actual needs of user research, avoiding excessive openness or restriction of permissions caused by single-dimensional authorization. While ensuring data security, it also ensures that legitimate users can obtain the data they need, balancing security and usability. Furthermore, by generating dynamic keys based on authorized structural domain fragments, conservatism scores, timestamps, and user-generated random numbers, the keys are tightly bound to data characteristics, user identity, and time. This dynamic association mechanism significantly improves the uniqueness and timeliness of the keys, reduces the risk of key reuse or forgery, and provides higher security for data transmission.

[0051] Then, the recombinant sequence is obtained by splicing authorized fragments, dynamic keys, and domain boundaries. The positional information of the domain boundaries ensures the accuracy of fragment splicing, making the recombinant sequence structurally close to the arrangement of the original sequence. Simultaneously, the decryption and verification process of the dynamic key further ensures the authenticity and integrity of the fragments used for splicing. Finally, by verifying the data integrity of the domain fragments and calculating the homology score between the recombinant and original sequences, and using a dual-condition determination of the recombinant sequence's validity, the integrity verification ensures that the data has not been tampered with or lost during transmission and splicing, and the homology score verifies the biological consistency between the recombinant and original sequences. This ensures that the sequence ultimately obtained by the user is not only formally complete but also possesses genuine biological functional characteristics, providing reliable data support for subsequent macromolecular analysis research. In summary, the macromolecular analysis data sharing management method provided in this embodiment effectively solves the significant shortcomings of current macromolecular analysis data sharing management methods in terms of the biological relevance of data preprocessing, the precision of access control, and the completeness of validity verification, failing to meet the needs of life science research for high-quality data sharing.

[0052] According to an embodiment of the present invention, a method for sharing and managing macromolecular analysis data is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0053] This embodiment provides a method for sharing and managing macromolecular analysis data. Figure 1 This is a flowchart of a macromolecular analysis data sharing and management method according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps:

[0054] Step S101: Obtain the original sequence data of the macromolecule and clean it, then perform a standardization transformation to obtain a standardized sequence.

[0055] Specifically, step S101 includes:

[0056] Step S1011: Set the window size of the sliding window to 5 residues. When the proportion of N bases in the sequence region of the original sequence data within the window exceeds a preset threshold, mark the above sequence region as a low-quality region and record the region coordinates.

[0057] Furthermore, the original macromolecular sequence is traversed sequentially in groups of five consecutive residues. At each window reached, the proportion of N bases is calculated. If the proportion exceeds a pre-set threshold (e.g., 15%-20% for next-generation sequencing), the corresponding sequence is marked as a "low-quality region," and the start and end coordinates of this region within the entire sequence are recorded for subsequent data processing to facilitate identification or avoidance. The functional importance and evolutionary conservation of different residues in macromolecules vary greatly; directly treating the sequence equally would result in the loss of crucial biological information. By calculating weights, subsequent analysis prioritizes functionally critical and evolutionarily stable residues, ensuring that data processing aligns with the true biological value of macromolecules and providing a basis for accurately identifying highly sensitive domains and rationally allocating access permissions.

[0058] Step S1012: Based on the baseline weights and conservation scores of the residues in the original sequence data, the residue weights are calculated using the following formula:

[0059]

[0060] Among them, For residues Biological weights, As the base value for weights, The conservation score of residues in homologous sequences,

[0061] Furthermore, residue weighting is a core step in data preprocessing. Its significance lies in quantifying the biological characteristics of macromolecules, such as functional importance, conservation, and sequencing quality, into calculable values. This provides a biological basis for subsequent fragmentation, licensing, and validation. The biological value of different residues in macromolecular sequences varies significantly; for example, the importance of amino acids in active sites differs drastically from that in non-functional regions. Weighting transforms abstract characteristics such as functional annotation, conservation, and sequencing quality into concrete weight values. For instance, critical functional residues are assigned high weights of 0.8-1.0, while residues in non-conserved regions are assigned low weights of 0.1-0.4, enabling subsequent algorithms to distinguish the biological importance of residues through numerical calculations.

[0062] Step S1013: Based on residue weights and a normalization formula, the original sequence is converted into a normalized sequence. The normalization formula is as follows:

[0063]

[0064] in, For the first , For residues The position index in the sequence, where L is the length of the original sequence. For residues Biological weights, This is the quality correction factor.

[0065] Furthermore, the core of standardization is to eliminate data format differences, but simply normalizing by position cannot reflect the differences in the biological value of residues. By introducing weights and quality correction factors, the standardization formula allows high-value residues to occupy higher weights in the standardized sequence. For example, residues at the same position, if they are active sites and have high sequencing quality, will have significantly higher standardized values ​​than non-functional regions and low-quality residues, ensuring that high-weight regions are preferentially protected during subsequent fragmentation.

[0066] Step S102: Use a domain identification tool to determine the domain boundaries of the macromolecules. Based on the domain boundaries and secondary structure prediction results, divide the standardized sequence into several domain segments and several flexible region segments.

[0067] Specifically, step S102 includes:

[0068] Step S1021: Use the HMMER tool to compare with the Pfam database to identify the domain boundaries in the normalized sequences. , ,in, Indicates the starting position of the m-th structural domain. and termination Location, for satisfying Annotate known structural domains within the structural domains.

[0069] Furthermore, specialized domain identification tools are used to compare the standardized macromolecular sequences with a database containing a large amount of known domain information. During the comparison process, the tool calculates the statistical significance index of the matching results, i.e., the E-value. When the statistical significance index is less than or equal to a specific threshold, such as 1e-5, this threshold is a standard recognized in the field that can effectively screen reliable domain matches, indicating that the corresponding known domain has been found, and determining its start and end positions in the standardized sequence, forming a set of domain boundaries. At the same time, based on the information stored in the database, these domains that pass the significance test are annotated with their functions, types, etc., clarifying their biological attributes. The domain identification tool HMMER is a mature software for domain analysis in the field of bioinformatics; the comparison database Pfam is a public resource that gathers a large amount of thoroughly studied domain data, accumulated and maintained by research institutions in the field over a long period of time; the E-value threshold is an empirical value determined in this field based on extensive experimental verification and statistical analysis, which can achieve a balance between "identifying true domains" and "filtering false positive matches". Domains in macromolecules are the core units that perform specific biological functions, and known domains have clear functional annotations and research foundations. By identifying and annotating known domains, we can quickly locate regions in standardized sequences that are "functionally clear and well-studied," providing direct biological evidence for subsequent functional sensitivity analysis and access control, and making the protection and authorization of high-value domains in data sharing more targeted.

[0070] Step S1022: For unannotated regions, predict secondary structures using PSIPRED. When the number of consecutive identical secondary structure residues is ≥8, mark the corresponding unannotated region as a potential domain. The resulting structural domain fragments are obtained by partitioning:

[0071]

[0072] in, For the standardized sequence fragments corresponding to the structural domains, This is biometric metadata, which includes domain IDs, secondary structure proportions, and functional tags for GO annotations.

[0073] Furthermore, for the sequence regions where no known domains were identified in step S1021, i.e., unannotated regions, secondary structure prediction tools are used to analyze secondary structures, such as α-helices and β-sheets. When the number of residues with the same consecutive secondary structure reaches or exceeds 8, the region is determined to have the potential to form a domain, is marked as a potential domain, and its start and end positions are determined. Subsequently, based on these marked domain boundaries, the corresponding sequence fragments are extracted from the normalized sequence, and the biometric metadata of the domain is compiled, including the domain ID for unique identification, the proportion of various secondary structures in the domain, and functional tags obtained from the gene ontology GO annotation, which together constitute the complete information of the domain fragment. PSIPRED, a secondary structure prediction tool, is a mature software widely used in bioinformatics and has been validated through extensive experiments. The standard of "≥8 consecutive identical secondary structure residues" is based on research into the smallest structural unit for domain formation, combined with experience in identifying potential domains within the field, and can effectively identify regions with domain formation potential. The domain ID in the biometric metadata is a unique identifier automatically generated by the system, the proportion of secondary structures is calculated statistically from the prediction results, and the GO annotation function tags come from public gene function annotation databases, continuously maintained and updated by researchers in the field. Potential domains not included in existing databases exist in macromolecular sequences; these regions may possess unique functions or research value. By predicting and labeling potential domains, we can, on the one hand, supplement the completeness of domain identification, avoiding the omission of valuable functional regions; on the other hand, by assigning biometric metadata to these potential domains, they can participate in subsequent analysis and authorization processes like known domains, ensuring data sharing covers all possible functional units of macromolecules.

[0074] Step S1023: Based on the flexibility index, splitting coefficient, and segment length, several flexible region segments are obtained. The flexibility index... for: ,in, The hydrophobicity index of residues, the resolution coefficient for: Fragment length for: , where s and e are the starting and ending positions of the flexible region, respectively.

[0075] Furthermore, the flexibility index is obtained by calculating the hydrophobicity of residues within the sequence region. The hydrophobicity index of each residue within the region is calculated, subtracted from 1, summed, and then divided by the region length. This reflects the flexibility of the region; regions with lower hydrophobicity are generally more flexible. The splitting coefficient is calculated based on the flexibility index using the formula: flexibility index multiplied by 5 plus 1, determining the basic number of segments to be split. Then, the fragment length is calculated by combining the start and end positions of the flexible region, i.e., the flexible region length divided by the splitting coefficient and rounded down. Finally, according to the calculated fragment lengths, the flexible region is split into several fragments, each retaining its corresponding sequence data and related characteristic information. The residue hydrophobicity index is an inherent parameter reflecting the physicochemical properties of residues, determined by research in the field of biochemistry, and can be obtained from authoritative biochemical databases or professional literature. The calculation formulas for the flexibility index, splitting coefficient, and fragment length are derived and designed based on the biological characteristics of the flexible region in this invention, transforming the qualitative description of flexibility into an operable splitting basis through mathematical operations. While flexible regions of macromolecules lack the defined, fixed functions of structural domains, they are crucial for conformational changes and functional regulation. Breaking flexible regions into segments allows for more detailed analysis of macromolecular dynamics, enabling subsequent research to be conducted based on these segments. Furthermore, these segmented flexible region fragments can collaborate with structural domain fragments in data sharing processes, ensuring the complete utilization of the macromolecular sequence characteristics and preventing the flexible regions from being overlooked or improperly processed due to the lack of defined structural domains.

[0076] Step S103: Obtain the functional sensitivity index, conservation score, functional importance score, and pathway correlation score of each domain segment, and calculate the functional sensitivity index of the current domain segment.

[0077] Specifically, the calculation formula is as follows:

[0078]

[0079] in, Let be the functional sensitivity index of the m-th structural domain segment. Let m be the conservatism score of the m-th structural domain segment. The functional importance score for the m-th structural domain segment. Let α+β+γ be the pathway connectivity score for the m-th domain segment, where α+β+γ=1.

[0080] Furthermore, three key biological scores were obtained for each domain fragment. For the conservation score, the domain fragment sequence was compared with a homology database to statistically analyze the degree of residue retention and consistency during evolution, resulting in a conservation score. A higher score indicates greater stability and more fundamental and critical function of the fragment during evolution. The functional importance score, based on the Gene Ontology GO annotation system, analyzed the core nature of the biological processes and molecular functions involved by the domain fragment, assigning values ​​through functional enrichment analysis and key functional pathway participation. A higher proportion of core functions resulted in a higher score. Then, the pathway association score was calculated by querying a pathway database to statistically analyze the number and criticality of disease-related and metabolic pathways involved by the domain fragment. More involvement in core and disease pathways resulted in a higher score. Finally, according to a pre-defined weighting rule, α, β, and γ corresponded to the weights of the three scores, and the sum of the three was 1. The three scores were then weighted and summed to obtain the functional sensitivity index of the domain fragment. The data foundation for the conservation score is the homology sequence database, composed of species genome and protein sequence data accumulated by research institutions over a long period of time. The functional importance score relies on the GO annotation database, which is continuously maintained and updated by research teams in the field through experimental verification and literature mining. The pathway association score is based on pathway databases such as KEGG, which gathers experimentally confirmed biological pathway information. The determination of the weights α, β, and γ requires combining the experience of field experts, a large amount of test data, and optimization after multiple sets of experimental comparisons and data analysis to ensure that they can reasonably reflect the contribution of the three dimensions to "functional sensitivity". The functional value and safety risks of macromolecular domain fragments vary significantly, and traditional single-dimensional assessment cannot comprehensively measure their sensitivity. By integrating the three dimensions of conservation, functional importance, and pathway association, the constructed functional sensitivity index can more accurately and comprehensively characterize the "value and risk level" of domain fragments. This index serves as the core basis for subsequent access authorization, allowing the data sharing process to set stricter access conditions for highly sensitive fragments, protecting core biological information while providing reasonable access for scientific research collaboration, thus resolving the contradiction between data security and sharing efficiency.

[0081] Step S104: Obtain the current user's research domain keyword set, functional tags of each structural domain segment, semantic similarity, and total number of functional tags, and calculate the user domain matching degree of the current structural domain segment.

[0082] Specifically, the calculation formula is as follows:

[0083]

[0084] in, For user domain matching degree, This is a collection of keywords for current user research. This refers to the t-th functional label of the m-th structural domain segment; Let m be the semantic similarity of the m-th structural domain segment, based on The ontology calculation shows that T is the total number of functional tags for the m-th structural domain segment.

[0085] Furthermore, a set of keywords related to the current user's research field is collected. These keywords can be submitted by the user during system registration, project application, or research direction submission, or extracted from the user's associated research papers and project documents through text mining. These keywords cover core content such as diseases, molecular mechanisms, and technical methods focused on by the user's research. Functional tags for each domain fragment are extracted. These functional tags are derived from bioinformatics databases such as Gene Ontology GO annotation and KEGG pathway annotation, providing standardized descriptions of the biological processes, molecular functions, and cellular components involved in the domain. Next, semantic similarity is calculated using the Unified Medical Language System (UMLS) ontology. UMLS integrates a vast amount of medical and biological terminology and semantic relationships. User keywords and domain functional tags are input, and the degree of semantic association between the two is calculated based on hierarchical relationships, synonym mapping, and semantic distance between terms. Finally, the total number of functional tags for the domain fragment is counted. The semantic similarity between each functional tag and the user keyword is summed and divided by the total number of tags to obtain the user domain matching degree, thus quantifying the degree of association between "user research needs - domain functional value". The keyword set for user research domains is sourced from user-submitted information and research text data, processed using natural language processing (NLP) techniques. Functional tags for domain segments rely on authoritative bioinformatics databases such as GO and KEGG, maintained by international research teams and continuously updated through experimental verification and literature review. Semantic similarity is calculated based on the UMLS ontology, which aggregates standardized terminology and semantic relationships in the medical and biological fields, constructed and maintained by the National Library of Medicine (NLM). The total number of functional tags is a direct count of the functional tags for individual domain segments, dynamically updated as the functional annotations of the domains are improved. Traditional data sharing permission determination often overlooks the compatibility between users' actual research needs and data functions, resulting in "users receiving irrelevant data and insufficient access to high-value data for users with specific needs." By calculating user domain matching, this approach accurately measures the alignment between the functions of domain segments and users' research directions from a semantic association perspective. Highly matched segments are more likely to correspond to users' research needs and should be given priority when granting authorization; segments with low matching scores can have access restricted, which not only avoids users obtaining redundant information, but also reduces the risk of leakage of core data due to broad authorization, allowing data sharing to shift from "extensive openness" to "precise flow on demand", thereby improving the efficiency of scientific research collaboration and the level of data security.

[0086] Step S105: If the functional sensitivity index and user domain matching degree of the current structural domain segment meet the first preset condition, grant the current user access rights to the current structural domain segment.

[0087] Specifically, the first preset condition is:

[0088]

[0089] in, This is the permission threshold.

[0090] Furthermore, the calculated functional sensitivity index of the current domain segment and the domain matching degree between the current user and the segment are obtained. Substituting these into a preset mathematical formula, the matching degree is multiplied by (1 minus the functional sensitivity index) to obtain a comprehensive evaluation value. This value is then compared with a pre-set permission threshold. If the comprehensive evaluation value is greater than or equal to the threshold, the authorization condition is met, and the user is granted access to the domain segment; if it is less than the threshold, access is denied, restricting data flow. The determination of the permission threshold requires extensive experimental simulations and real-world application scenario testing. It is derived through multiple rounds of adjustments and optimizations, analyzing the tolerance of different research fields for data sensitivity and the security and efficiency balance requirements of typical authorization cases. For example, for highly sensitive target data in drug development, the threshold can be set more strictly; for general domains shared in basic research, the threshold can be appropriately relaxed to ensure both the protection of core data and the avoidance of hindering reasonable scientific collaboration. Relying solely on data sensitivity or user identity can either lead to excessive restrictions resulting in low research efficiency or excessively broad authorization leading to data leakage risks. This step achieves a dynamic balance in two dimensions by constructing the condition "matching degree × (1 - sensitivity index) ≥ threshold": when the functional sensitivity of the domain fragment is high, the matching degree of the user's research needs to be high enough to offset the security risks brought about by high sensitivity; when the fragment sensitivity is low, the requirement for matching degree can be appropriately reduced, flexibly adapting to different scenarios. This mechanism makes authorization decisions more aligned with data value and scientific research needs, providing precise access control logic for the secure and efficient sharing of macromolecular analysis data, and ensuring that data circulates within a reasonable scope.

[0091] Step S106: Generate a dynamic key based on the authorized current structural domain fragment, conservatism score, timestamp, and user random number.

[0092] Specifically, the formula is as follows:

[0093]

[0094] in, For the current user Visit the Dynamic keys for each structural domain segment For the first A standardized sequence fragment of a structural domain segment, Let m be the conservatism score of the m-th structural domain segment. Millisecond-level timestamps For the current user A unique random number is generated during registration. For hash functions, This is for string concatenation operations.

[0095] Furthermore, multiple key pieces of information from the authorized structural domain fragments are integrated. This includes obtaining the standardized sequence fragment and conservatism score of the fragment; generating a millisecond-level timestamp using the system clock, accurate to the millisecond to ensure uniqueness in the time dimension; and extracting a unique random number generated during user registration, unique to each user, generated and associated with the user account using a random number algorithm during the registration process. These information are then concatenated in a fixed order to form a composite string containing biometrics, time, and user identifier. Finally, this composite string is input into the SHA256 hash function to generate a fixed-length, irreversible hash value, i.e., the dynamic key. Traditional static keys are easily stolen and reused, posing a risk of data forgery. This step uses a dynamic key mechanism to ensure that each authorized access action corresponds to a unique key: the biometric information of the structural domain fragment ensures that the key is bound to the data function; the timestamp allows the key to change dynamically over time; and the user-specific random number strengthens the user's uniqueness of the key. The one-way nature of the hash function ensures that the key cannot be reverse-engineered from the ciphertext, making it difficult to crack even if intercepted during transmission. This mechanism enhances security from the source of key generation, builds a solid security defense for subsequent data transmission and recombination verification, and ensures that authorized access to structural domain fragments is in a state of encrypted protection throughout the entire process, which meets the high-value and highly sensitive security requirements of macromolecular analysis data.

[0096] Step S107: The structural domain fragments are spliced ​​together according to their positional order in the original sequence data to obtain the recombinant sequence.

[0097] Specifically, authorized fragments are sorted based on the natural positional order of domains in the original sequence data—arranging the fragments sequentially according to the start and end positions of the domain boundary records to ensure that the splicing order is consistent with the structural order of the original macromolecular sequence. The function of macromolecular sequences depends on the ordered arrangement of domains. In traditional data sharing, if scattered fragments are directly provided, users must manually splice them, and the authenticity of the data cannot be guaranteed. By splicing according to the original positional order, the natural structural relationship of the macromolecular sequence is restored, ensuring that the recombinant sequence has biological analytical value. Dynamic key verification is introduced to prevent unauthorized or tampered fragments from entering the splicing process, ensuring the authenticity and legitimacy of the recombinant sequence data. This mechanism ensures that the recombinant sequences obtained by users are not only complete functional sequences that meet research needs but also trusted data that has undergone authorization and security verification.

[0098] Step S108: Verify the data integrity of the recombinant sequence, calculate the homology score between the recombinant sequence and the original sequence. If the data integrity verification result and the homology score meet the second preset condition, the recombinant sequence is determined to be valid, and the recombinant sequence is sent to the current user's client.

[0099] Specifically, the second pre-defined condition is to determine the validity of the recombinant sequence by combining the data integrity verification result and the homology score. The process is as follows: First, the data integrity verification is checked. If the checksums of all domain fragments are consistent with the original stored checksums, the data integrity condition is met. Next, the homology score is checked, comparing the calculated homology score with a pre-set threshold. This threshold is determined based on industry standards for macromolecular sequence analysis, extensive experimental data, and the requirements for sequence validity in different application scenarios. For example, in macromolecular analysis related to drug development, the threshold is set higher to ensure high sequence accuracy. If the score is greater than or equal to this threshold, the homology condition is met. When both data integrity and homology score meet their respective conditions, the recombinant sequence is deemed valid and sent to the user's client for subsequent scientific analysis. If either condition is not met, the recombinant sequence is deemed invalid and will not be sent to the user. The anomaly is also recorded for troubleshooting. We rigorously control the quality of the data delivered to users from two dimensions: data integrity and biological validity. This ensures that the recombinant sequences obtained by users are free from data flaws and possess reliable biological analytical value, providing solid data support for functional studies, disease mechanism exploration, and other work based on these sequences, and preventing invalid data from affecting the research process and results.

[0100] Specifically, the data integrity verification formula is as follows:

[0101]

[0102] in, Let m be the checksum of the m-th structural domain segment. Let i be the normalized value of the i-th residue. , Let be the start and end positions of the m-th structural domain segment.

[0103] Furthermore, the data integrity verification step primarily aims to ensure that the data of each domain fragment in the recombinant sequence has not been tampered with or lost during transmission, splicing, or other processes. For the m-th domain fragment, the normalized value of each residue is obtained sequentially from its start position to its end position. Then, the normalized value of each residue is multiplied by its position index i within the fragment. These products are summed, and finally, the result is taken modulo 2^32 to obtain the checksum of the domain fragment. The calculated checksum is compared with the checksum pre-stored during the fragment licensing process. If they match, the data is intact and has not been tampered with; if they do not match, the data may be corrupted or maliciously modified. The purpose of this step is to ensure the authenticity of the domain fragments at the data level, laying a solid foundation for the subsequent recombinant sequence to accurately reflect the characteristics of the original macromolecular sequence and avoiding erroneous conclusions in subsequent analyses due to incomplete or tampered data.

[0104] Specifically, the formula for calculating the homology score is as follows:

[0105]

[0106] Where H is the homology score, with a value ranging from 0 to 1. This represents the number of matching residues between the recombinant sequence and the original sequence. The length of the original sequence. This represents the length of the recombinant sequence.

[0107] Furthermore, homology score calculation verifies the similarity between the recombinant sequence and the original sequence from a biological perspective. The process employs a global alignment algorithm to compare the recombinant sequence and the original sequence residue-by-residue, counting the number of matching residues. Here, the original sequence is the macromolecular original sequence obtained at the beginning of the process, and the recombinant sequence is the sequence assembled along domain boundaries in step S107. According to the formula, twice the number of matching residues is divided by the sum of the lengths of the original and recombinant sequences to obtain the homology score. This score ranges from 0 to 1; a higher score indicates a higher consistency between the recombinant and original sequences and a stronger biological correlation. The aim is to ensure that the recombinant sequence is not only complete but also highly consistent with the original sequence in terms of biological characteristics, possessing research analysis value. This prevents situations where the sequence is correctly assembled but its function deviates completely from the original due to incorrect fragment combinations, ensuring that the recombinant sequence received by the user truly reflects the biological information of the macromolecule.

[0108] The macromolecule analysis data sharing and management method provided in this embodiment firstly cleans the original macromolecule sequence data and calculates residue weights by combining the base weights of residues and conservation scores, then converts it into a standardized sequence. This process not only removes interference information from low-quality regions in the original data, improving data reliability, but also incorporates the biological characteristics of the sequence into the standardization process through residue weights, ensuring that the standardized sequence not only has data consistency but also retains key biological information. Secondly, a domain identification tool is used to determine domain boundaries, and the results of secondary structure prediction are combined to divide the data into domain fragments and flexible region fragments. Each fragment contains sequence data and biological metadata, ensuring that the divided fragments conform to the natural structure and functional characteristics of the macromolecule. The integrity of the domain fragments is maintained, avoiding functional unit destruction caused by random segmentation, while the biological metadata provides direct biological basis for subsequent permission determination and data analysis. Then, by acquiring the functional sensitivity index, conservation score, functional importance score, and pathway association score of the domain fragments and calculating the functional sensitivity index, the biological functional sensitivity of each domain fragment is quantified, clarifying its importance in macromolecular function execution. This allows for more stringent protection of highly sensitive and important fragments, reducing the risk of leakage of core biological information. Furthermore, by combining the user's research domain keyword set with the functional tags of the domain fragments, and using semantic similarity and the total number of functional tags to calculate the user domain matching degree, the relevance between user needs and fragment functions can be accurately measured. This ensures that the fragments obtained by users are highly relevant to their research direction, reducing the transmission and access of irrelevant data and improving the targeting and efficiency of data sharing. Finally, by combining the functional sensitivity index and user domain matching degree as authorization conditions, a two-dimensional access control system is implemented. This considers both the inherent sensitivity of the domain fragments and the actual needs of user research, avoiding excessive openness or restriction of permissions caused by single-dimensional authorization. While ensuring data security, it also ensures that legitimate users can obtain the data they need, balancing security and usability. Furthermore, by generating dynamic keys based on authorized structural domain fragments, conservatism scores, timestamps, and user-generated random numbers, the keys are tightly bound to data characteristics, user identity, and time. This dynamic association mechanism significantly improves the uniqueness and timeliness of the keys, reduces the risk of key reuse or forgery, and provides higher security for data transmission.

[0109] Then, the recombinant sequence is obtained by splicing authorized fragments, dynamic keys, and domain boundaries. The positional information of the domain boundaries ensures the accuracy of fragment splicing, making the recombinant sequence structurally close to the arrangement of the original sequence. Simultaneously, the decryption and verification process of the dynamic key further ensures the authenticity and integrity of the fragments used for splicing. Finally, by verifying the data integrity of the domain fragments and calculating the homology score between the recombinant and original sequences, and using a dual-condition determination of the recombinant sequence's validity, the integrity verification ensures that the data has not been tampered with or lost during transmission and splicing, and the homology score verifies the biological consistency between the recombinant and original sequences. This ensures that the sequence ultimately obtained by the user is not only formally complete but also possesses genuine biological functional characteristics, providing reliable data support for subsequent macromolecular analysis research. In summary, the macromolecular analysis data sharing management method provided in this embodiment effectively solves the significant shortcomings of current macromolecular analysis data sharing management methods in terms of the biological relevance of data preprocessing, the precision of access control, and the completeness of validity verification, failing to meet the needs of life science research for high-quality data sharing.

[0110] The above are embodiments of the macromolecular analysis data sharing management system provided in this application. Other embodiments of the macromolecular analysis data sharing management system provided in this application will be described below. Please refer to the following for details.

[0111] This embodiment also provides a macromolecular analysis data sharing and management system, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the systems described in the following embodiments are preferably implemented in software, hardware implementations, or combinations of software and hardware, are also possible and contemplated.

[0112] This embodiment provides a macromolecular analysis data sharing and management system, such as Figure 2 As shown, it includes:

[0113] The standardization module 201 is used to acquire the raw sequence data of macromolecules, clean them, and perform standardization transformation to obtain standardized sequences.

[0114] The segmentation module 202 is used to determine the domain boundaries of the above-mentioned macromolecules using a domain identification tool, and to divide the above-mentioned standardized sequence into several domain segments and several flexible region segments based on the domain boundaries and secondary structure prediction results.

[0115] The first calculation module 203 is used to obtain the functional sensitivity index, conservation score, functional importance score and pathway correlation score of each domain segment, and calculate the functional sensitivity index of the current domain segment.

[0116] The second calculation module 204 is used to obtain the current user's research domain keyword set, functional tags of each structural domain segment, semantic similarity and total number of functional tags, and to calculate the user domain matching degree of the current structural domain segment.

[0117] Access module 205 is used to grant the current user access rights to the current structural domain segment if the functional sensitivity index and user domain matching degree of the current structural domain segment meet the first preset condition.

[0118] The generation module 206 is used to generate a dynamic key based on the authorized current domain fragment, the conservatism score, the timestamp, and the user random number;

[0119] The splicing module 207 is used to splice the fragments of each structural domain according to their positional order in the original sequence data to obtain the recombined sequence;

[0120] The shared module 208 is used to verify the data integrity of the recombined sequence, calculate the homology score between the recombined sequence and the original sequence, and if the data integrity verification result and the homology score meet the second preset condition, the recombined sequence is determined to be valid and the recombined sequence is sent to the current user's client.

[0121] The macromolecule analysis data sharing management system provided in this embodiment first cleans the original macromolecule sequence data, calculates residue weights by combining the base weights of residues and conservation scores, and then converts it into a standardized sequence. This process not only removes interference information from low-quality regions in the original data, improving data reliability, but also incorporates the biological characteristics of the sequence into the standardization process through residue weights, ensuring that the standardized sequence not only has data consistency but also retains key biological information. Secondly, it uses a domain identification tool to determine domain boundaries and combines secondary structure prediction results to divide domain fragments and flexible region fragments. Each fragment contains sequence data and biological metadata, ensuring that the divided fragments conform to the natural structure and functional characteristics of the macromolecule. The integrity of the domain fragments is maintained, avoiding functional unit destruction caused by random segmentation, while the biological metadata provides direct biological basis for subsequent permission determination and data analysis. Then, by acquiring the functional sensitivity index, conservation score, functional importance score, and pathway association score of the domain fragments and calculating the functional sensitivity index, the biological functional sensitivity of each domain fragment is quantified, clarifying its importance in macromolecular function execution. This allows for more stringent protection of highly sensitive and important fragments, reducing the risk of leakage of core biological information. Furthermore, by combining the user's research domain keyword set with the functional tags of the domain fragments, and using semantic similarity and the total number of functional tags to calculate the user domain matching degree, the relevance between user needs and fragment functions can be accurately measured. This ensures that the fragments obtained by users are highly relevant to their research direction, reducing the transmission and access of irrelevant data and improving the targeting and efficiency of data sharing. Finally, by combining the functional sensitivity index and user domain matching degree as authorization conditions, a two-dimensional access control system is implemented. This considers both the inherent sensitivity of the domain fragments and the actual needs of user research, avoiding excessive openness or restriction of permissions caused by single-dimensional authorization. While ensuring data security, it also ensures that legitimate users can obtain the data they need, balancing security and usability. Furthermore, by generating a dynamic key based on authorized structural domain fragments, conservatism scores, timestamps, and user-generated random numbers, the key is tightly bound to data characteristics, user identity, and time. This dynamic association mechanism significantly improves the uniqueness and timeliness of the key, reduces the risk of key reuse or forgery, and provides higher security for data transmission. Then, the reconstructed sequence is obtained by splicing the authorized fragments, the dynamic key, and the structural domain boundaries. The positional information of the structural domain boundaries ensures the accuracy of fragment splicing, making the reconstructed sequence structurally close to the arrangement of the original sequence. Simultaneously, the decryption and verification process of the dynamic key further ensures the authenticity and integrity of the fragments used for splicing.Finally, by verifying the data integrity of the domain fragments and calculating the homology score between the recombinant sequence and the original sequence, and using a dual-condition determination of the validity of the recombinant sequence, the system ensures that the data has not been tampered with or lost during transmission and assembly through integrity verification, and verifies the biological consistency between the recombinant sequence and the original sequence through homology score verification. This ensures that the sequence ultimately obtained by the user is not only formally complete but also possesses genuine biological functional characteristics, providing reliable data support for subsequent macromolecular analysis research. In summary, the macromolecular analysis data sharing management system provided in this embodiment effectively solves the significant shortcomings of current macromolecular analysis data sharing management systems in terms of the biological relevance of data preprocessing, the precision of access control, and the completeness of validity verification, failing to meet the needs of life science research for high-quality data sharing.

[0122] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

Claims

1. A method for sharing and managing macromolecular analysis data, characterized in that, The method includes: The raw sequence data of macromolecules are obtained and cleaned, and then standardized to obtain standardized sequences. The domain boundaries of the macromolecule are determined using a domain identification tool. Based on the domain boundaries and secondary structure prediction results, the normalized sequence is divided into several domain segments and several flexible region segments. The functional sensitivity index of the current structural domain segment is calculated using the following formula: ; in, Let be the functional sensitivity index of the m-th structural domain segment. Let m be the conservatism score of the m-th structural domain segment. The functional importance score for the m-th structural domain segment. The pathway connectivity score for the m-th domain segment is α+β+γ=1; Obtain the current user's research domain keyword set, functional tags for each structural domain segment, semantic similarity, and total number of functional tags. Calculate the user domain matching degree for the current structural domain segment using the following formula: ; in, For user domain matching degree, This is a collection of keywords for current user research. This refers to the t-th functional label of the m-th structural domain segment; Let m be the semantic similarity of the m-th structural domain segment, based on The ontology calculation shows that T is the total number of functional tags for the m-th structural domain segment; If the functional sensitivity index and user domain matching degree of the current structural domain segment meet the first preset condition, the current user is granted access to the current structural domain segment. Generate a dynamic key based on the authorized current domain fragment, conservatism score, timestamp, and user-generated random number; The structural domain segments are spliced ​​together according to their positional order in the original sequence data to obtain the recombined sequence; Verify the data integrity of the recombinant sequence, calculate the homology score between the recombinant sequence and the original sequence, and if the data integrity verification result and the homology score meet the second preset condition, determine that the recombinant sequence is valid and send the recombinant sequence to the current user's client.

2. The method according to claim 1, characterized in that, The process of acquiring and cleaning the raw sequence data of macromolecules, and then performing a standardization transformation to obtain a standardized sequence includes: The sliding window size is set to 5 residues. When the proportion of N bases in the sequence region within the window exceeds a preset threshold, the sequence region is marked as a low-quality region and the region coordinates are recorded. The residue weights are calculated based on the baseline weights and conservation scores of the residues in the original sequence data, using the following formula: ; Among them, For residues Biological weights, As the base value for weights, ; Score the conservation of residues in homologous sequences; The original sequence is converted into a normalized sequence based on residue weights and a normalization formula, which is as follows: ; in, Let i be the normalized value of the i-th residue. For residues The position index in the sequence, where L is the length of the original sequence. For residues Biological weights, This is the quality correction factor.

3. The method according to claim 1, characterized in that, The process involves using domain identification tools to determine the domain boundaries of the macromolecule, and based on the domain boundaries and secondary structure prediction results, dividing the normalized sequence into several domain segments and several flexible region segments, including: The HMMER tool was used to compare the data against the Pfam database to identify domain boundaries in the normalized sequences. , ,in, Indicates the starting position of the m-th structural domain. and termination Location, for satisfying Annotate the known structural domains within the structural domains; For unannotated regions, secondary structures are predicted using PSIPRED. When the number of consecutive identical secondary structure residues is ≥8, the corresponding unannotated region is marked as a potential domain. ; The resulting structural domain fragments are obtained: ; in, For the standardized sequence fragments corresponding to the structural domains, The biometric metadata includes domain IDs, secondary structure proportions, and functional tags for GO annotations. Several flexible region segments are obtained by splitting based on the flexibility index, splitting coefficient, and segment length. The flexibility index... for: ,in, The hydrophobicity index of residues, the resolution coefficient for: Fragment length for: , where s and e are the starting and ending positions of the flexible region, respectively.

4. The method according to claim 3, characterized in that, The first preset condition is: ; in, This is the permission threshold.

5. The method according to claim 4, characterized in that, The dynamic key is generated based on the authorized current domain fragment, conservatism score, timestamp, and user random number, using the following formula: ; in, For the current user Visit the Dynamic keys for each structural domain segment For the first A standardized sequence fragment of a structural domain segment, Let m be the conservatism score of the m-th structural domain segment. Millisecond-level timestamps For the current user A unique random number is generated during registration. For hash functions, This is for string concatenation operations.

6. The method according to claim 5, characterized in that, The data integrity verification formula is as follows: ; in, Let m be the checksum of the m-th structural domain segment. Let i be the normalized value of the i-th residue. , Let be the start and end positions of the m-th structural domain segment.

7. The method according to claim 6, characterized in that, The formula for calculating the homology score is as follows: ; Where H is the homology score, with a value ranging from 0 to 1. This represents the number of matching residues between the recombinant sequence and the original sequence. The length of the original sequence. This represents the length of the recombinant sequence.

8. A macromolecular analysis data sharing and management system, characterized in that, The system includes: The standardization module is used to acquire and clean the raw sequence data of macromolecules, and then perform a standardization transformation to obtain standardized sequences. The segmentation module is used to determine the domain boundaries of the macromolecule using a domain identification tool, and to divide the normalized sequence into several domain segments and several flexible region segments based on the domain boundaries and secondary structure prediction results. The first calculation module is used to calculate the functional sensitivity index of the current structural domain segment. The calculation formula is as follows: ; in, Let be the functional sensitivity index of the m-th structural domain segment. Let m be the conservatism score of the m-th structural domain segment. The functional importance score for the m-th structural domain segment. The pathway connectivity score for the m-th domain segment is α+β+γ=1; The second calculation module is used to obtain the current user's research domain keyword set, functional tags of each structural domain segment, semantic similarity, and total number of functional tags, and to calculate the user domain matching degree of the current structural domain segment. The calculation formula is as follows: ; in, For user domain matching degree, This is a collection of keywords for current user research. This refers to the t-th functional label of the m-th structural domain segment; Let m be the semantic similarity of the m-th structural domain segment, based on The ontology calculation shows that T is the total number of functional tags for the m-th structural domain segment; The access module is used to grant the current user access rights to the current structural domain segment if the functional sensitivity index and user domain matching degree of the current structural domain segment meet the first preset condition. The generation module is used to generate dynamic keys based on the authorized current domain fragment, conservatism score, timestamp, and user random number; The splicing module is used to splice the fragments of each structural domain according to their positional order in the original sequence data to obtain the recombined sequence; The sharing module is used to verify the data integrity of the recombined sequence, calculate the homology score between the recombined sequence and the original sequence, and if the data integrity verification result and the homology score meet the second preset condition, the recombined sequence is determined to be valid and the recombined sequence is sent to the current user's client.

Citation Information

Patent Citations

  • Efficient sharing method for biomolecular data

    CN102411572A

  • Protein aided design and analysis system and method based on cloud computing platform

    CN120072027A