Scientific data evaluation method and system based on FAIR principle

By constructing a scientific data evaluation method based on the FAIR principle, quantifiable indicators and multi-source information verification are used to solve the problem of inconsistent scientific data evaluation standards, the automated evaluation and open sharing of scientific data are realized, and the availability of data is improved.

CN120471494APending Publication Date: 2025-08-12COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510365943.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

There is a lack of unified scientific data FAIR evaluation standards in the existing technology. Traditional evaluation methods rely on qualitative indicators and cannot meet the needs of efficient utilization of large-scale scientific data. There are subjectivity and generalization problems.

Method used

A scientific data evaluation method based on FAIR principle is constructed, quantifiable evaluation indicators are used, metadata information is obtained using the unique identifier of scientific data, and FAIR evaluation results are generated and improved suggestions are provided through multi-source information replenishment and verification.

Benefits of technology

The automated FAIR evaluation of scientific data is realized, which improves evaluation efficiency and impartiality, promotes the open sharing and availability of scientific data, and ensures that the data sets meet the standards of discoverability, accessibility, interoperability, reusability and manageability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471494A_ABST
    Figure CN120471494A_ABST
Patent Text Reader

Abstract

The invention discloses a scientific data evaluation method and system based on an FAIR principle, and the method comprises the steps: constructing a set of autonomous scientific data FAIR evaluation model, carrying out the comprehensive evaluation of the FAIR level, data compliance and data availability of a data set, and guaranteeing that the data set meets the evaluation indexes which can be found, obtained, interoperable, reusable and manageable; and obtaining the evaluation result of the availability of the scientific data according to the result of the FAIR evaluation index of the scientific data. Advanced data acquisition and extraction capabilities are combined, automation of a data evaluation process is realized, manual intervention is greatly reduced, and evaluation efficiency and fairness are improved; internationally open and domestic independent third-party data sources are adopted, so that multi-source information supplementation and cross check are realized, and the obtained evaluation result is objective and fair. According to the method, scientific data management efficiency can be improved, scientific data opening and sharing are promoted, and scientific data availability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of scientific data, and in particular, to a scientific data evaluation method and system based on the FAIR principle. Background Art

[0002] Scientific data is a crucial cornerstone of digital intelligence-driven innovation. With the development of data-intensive scientific research paradigms, the need for governance regarding the acquisition, accessibility, and interoperability of open data has become more pressing. Against this backdrop, the FAIR (Discoverable, Accessible, Interoperable, Reusable) principles for scientific data were proposed in 2016. These principles aim to define a set of rules and regulations that enable both machines and humans to discover, access, interoperate, and reuse data and metadata.

[0003] Since its inception, the FAIR principle has been widely applied in data governance across multiple fields. Data governance practices that adhere to the FAIR principle help generate high-quality AI-ready data. However, the current implementation of the FAIR principle suffers from subjectivity and generalization issues, and lacks unified evaluation standards. First, due to the diverse formats and complex structures of scientific research datasets, the data governance requirements of various disciplines vary, making it difficult to adopt unified FAIR evaluation standards for data in different fields. Second, traditional data evaluation methods rely primarily on qualitative indicators and lack quantitative means. The evaluation process is cumbersome and time-consuming, especially when dealing with large-scale datasets. Traditional evaluation methods have delayed feedback and cannot meet the needs of efficient use of scientific data. Summary of the Invention

[0004] To overcome the limitations of current scientific data FAIR capability assessment methods, this paper aims to develop an independent scientific data FAIR assessment method based on my country's policies and latest requirements for scientific data development. This method designs a set of quantifiable evaluation indicators based on the FAIR principle. It uses the unique identifier of scientific data to extract relevant metadata information. It then uses internationally available and domestically independent information sources to complete and verify multi-source information. Finally, it obtains FAIR assessment results and improvement suggestions for scientific data, promoting the open sharing of scientific data and improving its usability.

[0005] The technical solution adopted in the present invention is as follows:

[0006] A scientific data evaluation method based on the FAIR principle includes the following steps:

[0007] Building an evaluation index system for scientific data based on the FAIR principle. The evaluation index system includes the dimensions of discoverability, accessibility, interoperability, reusability, and manageability. Several evaluation indicators are defined for each dimension and stored in a database.

[0008] Collect scientific data to be evaluated, parse the unique identifiers of the scientific data to be evaluated, and generate standardized URLs;

[0009] Based on standardized URLs, obtain metadata information of data and perform multi-source metadata alignment and verification;

[0010] Based on the aligned and verified metadata information and the constructed evaluation indicator system, the evaluation indicators of each dimension are calculated;

[0011] Generate a feedback report based on the evaluation indicator calculation results.

[0012] Furthermore, the discoverability dimension includes the following evaluation indicators: using a globally unique identifier, and the identifier can be normally resolved and located to the data resource page; the metadata contains descriptive core elements, including creator, title, data identifier, publishing organization, publication date, abstract and keywords; the metadata accurately describes the access address, file name, data volume and data format of the entity data; the metadata is provided in a machine-readable format and can be machine-retrieved.

[0013] Furthermore, the accessibility dimension includes the following evaluation indicators: metadata supports user online access; data open sharing methods are clear, including access levels and access conditions; entity data supports user access.

[0014] Furthermore, the interoperability dimension includes the following indicators: data supports universal interoperability protocols.

[0015] Furthermore, the reusability dimension includes the following indicators: metadata describing the permission statement for data use; metadata clarifying the creation or generation source information of the data; and data complying with the standards of the relevant scientific research field.

[0016] Furthermore, the manageable dimensions include the following indicators: data complies with security and ethical review requirements; metadata complies with compliance requirements, and core metadata complies with domestic and international standards, such as the number of abstract words, number of keywords, subject classification, etc.; an independent identification system is used; data supports traceability, and metadata contains associated entity information and related projects.

[0017] Furthermore, the multi-source metadata information alignment and verification includes:

[0018] Cleaning: Remove irrelevant or redundant fields from the acquired metadata to ensure data simplicity;

[0019] Metadata mapping: Map the extracted metadata according to the preset standard metadata vocabulary, converting metadata information from different sources into a unified key-value pair format;

[0020] Merge: Merge metadata from different sources into a complete metadata collection. If a key already exists during the metadata merging process, the value type is checked. If the value is a string, a similarity algorithm is used to determine whether the existing value needs to be replaced. If the value is an array, the new value is merged with the existing array and deduplication is performed to ensure the uniqueness of the elements in the array.

[0021] Output: The metadata processed by the above steps will be output in the form of standardized key-value pairs.

[0022] Furthermore, the metadata information based on alignment and verification and the constructed evaluation index system calculates the evaluation index of each dimension, including:

[0023] According to the constructed FAIR evaluation index system, the scores of each indicator in the five dimensions are calculated as shown in the following formula:

[0024]

[0025] Among them, X total The full score evaluation rules set in each dimension, X∈{F,A,I,R,G}, i is the specific indicator item number of each dimension, is the scoring weight of the i-th indicator item, is the score of the i-th indicator item; thus, the score of each dimension can be obtained as shown in the following formula:

[0026]

[0027] Where n is the number of specific indicators in each dimension; the average value of all dimension scores is calculated according to the above formula, as shown below:

[0028]

[0029] Where S is the overall FAIR capability completion of scientific data.

[0030] Furthermore, the feedback report includes evaluation scores, visual reports and corresponding optimization suggestions.

[0031] A scientific data evaluation system based on the FAIR principle, comprising:

[0032] Construct an evaluation index system module, which is used to construct an evaluation index system for scientific data based on the FAIR principle. The evaluation index system includes the dimensions of discoverability, accessibility, interoperability, reusability, and manageability. Several evaluation indicators are defined for each dimension and stored in the database.

[0033] The data collection module is used to collect scientific data to be evaluated, parse the unique identifier of the scientific data to be evaluated, and generate a standardized URL;

[0034] The data processing module is used to obtain metadata information of the data based on the standardized URL and align and verify the multi-source metadata information;

[0035] The calculation and evaluation module is used to calculate the evaluation indicators of each dimension based on the metadata information of alignment and verification and the constructed evaluation indicator system;

[0036] The feedback module is used to generate a feedback report based on the evaluation indicator calculation results.

[0037] Based on the FAIR principles and in combination with my country's scientific data development policies and latest requirements, this paper constructs an independent scientific data FAIR evaluation method and system to comprehensively evaluate the FAIR level, data compliance, and data availability of datasets, ensuring that datasets meet the evaluation indicators of discoverability, accessibility, interoperability, reusability, and manageability, providing solid guarantees for the open sharing and reuse of data.

[0038] In summary, the scientific data evaluation method based on the FAIR principle proposed in this invention can automatically evaluate the FAIR compliance of scientific data and provide improvement suggestions, providing technical support for improving the standardized management level of data, promoting the open sharing of scientific data, and improving the availability of scientific data. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 This is a flow chart of a scientific data evaluation method based on the FAIR principle of the present invention.

[0040] Figure 2 yes Figure 1 Schematic diagram of constructing the FAIR indicator system in the method flowchart shown.

[0041] Figure 3 yes Figure 1 Schematic diagram of metadata acquisition and multi-source metadata information alignment and verification in the method flowchart shown.

[0042] Figure 4 yes Figure 1 Schematic diagram of the FAIR evaluation index calculation process in the method flowchart shown. DETAILED DESCRIPTION

[0043] The present invention will be further explained below with reference to the accompanying drawings using specific implementation methods.

[0044] The present invention relates to a scientific data evaluation method based on the FAIR principle. The following are the specific implementation steps, including the five steps of indicator construction evaluation indicator system, data collection, data processing, calculation evaluation and feedback.

[0045] 1. Build an evaluation indicator system

[0046] In this paper, we first need to build a scientific data evaluation index system based on the FAIR principle. The construction of this index system is the foundation of the scientific data evaluation method and has the following key steps:

[0047] 1.1 Definition of primary indicators: citing the four core dimensions of data discoverability, accessibility, interoperability, and reusability based on the FAIR principles. Taking into account the characteristics and requirements of my country in scientific data management, the dimension of "governable" has been added.

[0048] 1.2 Define specific evaluation indicators for each dimension: Design specific, quantifiable, and accessible evaluation indicators for each dimension.

[0049] 1.3 Storing evaluation indicators in the database: The independent scientific data FAIR indicator system is stored in the database for subsequent use.

[0050] 2. Data Collection

[0051] After building the FAIR evaluation index system, it is necessary to collect the scientific data to be evaluated, parse the data's unique identifiers, and generate standardized URLs. During the collection process, this information must be verified to ensure its validity. Data collection methods include the following:

[0052] 2.1 Collect data digital object identifier (DOI).

[0053] 2.2 Collect Chinese Scientific and Technological Resource Identifiers (CSTR).

[0054] 2.3 Collect Uniform Resource Locators (URLs).

[0055] 3. Data Processing

[0056] After completing the data collection in step 2 and obtaining the data standardized URL, the data metadata will be obtained and the metadata information will be aligned and verified.

[0057] 3.1 Metadata acquisition includes the following operations:

[0058] 3.1.1 Requesting Web Data: Use a tool like Selenium to request the standardized URL for the data collected in Step 2 to retrieve the webpage content associated with the identifier. Some webpages may require JavaScript to complete before retrieving data, so a delayed loading time is required to ensure complete metadata extraction.

[0059] 3.1.2 Extract metadata: Extract basic metadata from the page DOM elements.

[0060] 3.1.3 Call external API to obtain more metadata: Call external API to obtain more metadata.

[0061] 3.2 Multi-source metadata information completion and verification includes the following operations:

[0062] 3.2.1 Cleaning: Clean irrelevant or redundant fields to ensure data simplicity.

[0063] 3.2.2 Metadata mapping: Map the extracted metadata according to the preset standard metadata vocabulary, and convert metadata fields from different sources into consistent key-value pairs.

[0064] 3.2.3 Merge: Merge metadata from different sources into a complete metadata information. If it is found that some keys already exist during the metadata merging process, the similarity algorithm will be used to determine whether the existing values need to be replaced to ensure the real-time nature of the metadata.

[0065] 3.2.4 Output: The metadata processed through the above steps will be output in the form of standardized key-value pairs to provide data support for subsequent FAIR evaluation.

[0066] 4. Computational Evaluation

[0067] After completing the standardized information obtained in step 3, the FAIR capability will be evaluated. According to the FAIR evaluation index system constructed in step 1, the five indicators are calculated and evaluated respectively. The following steps are included:

[0068] 4.1 Loading evaluation rules: Read the FAIR evaluation system data constructed in step 1 from the database for subsequent calculation and evaluation.

[0069] 4.2 Evaluation score: Traverse the evaluation rules of the five dimensions and calculate the scores of each indicator in the five dimensions Where X∈{F,A,I,R,G}, i represents the number of each indicator item in the five dimensions.

[0070] 4.3 Calculate the average score: Calculate the average of all normalized scores obtained in the previous step as the FAIR capability completion of the data.

[0071] 5. Feedback

[0072] Based on the calculation and evaluation results obtained in step 4, a detailed feedback report is generated, which includes the following steps:

[0073] 5.1 Result feedback: Create a visual report based on the scores of the evaluation results obtained in step 4 to facilitate subsequent analysis.

[0074] 5.2 Improvement suggestions: Compare each indicator in the calculated evaluation results obtained in step 4 with its total score, and provide relevant improvement suggestions.

[0075] Figure 1 A flow chart of a scientific data evaluation method based on the FAIR principle is provided as a flowchart of an embodiment of the present invention. Figure 1 As shown in FIG, the present invention mainly includes 5 parts. The first part is as shown in FIG. Figure 2 As shown in the figure, the evaluation index system is constructed based on the FAIR principle and stored in the database for subsequent use; the second part is data collection, which parses the scientific data to be evaluated to obtain the unique identification information of the data and forms a standardized URL; the third part is as shown in the figure. Figure 3 As shown in the figure, the information obtained in the second part is processed to obtain multi-source metadata information and to complete and verify it to ensure the uniformity of the data; the fourth part is as shown in the figure. Figure 4 As shown in Figure 1, the information obtained in the previous step is calculated and evaluated according to the indicator system. The fifth part is feedback, which generates evaluation feedback and improvement suggestions based on the results of the previous step. The following describes the detailed implementation of each part.

[0076] 1. Build an evaluation indicator system

[0077] 1.1 Implemented four indicators and requirements in indicator F (discoverability):

[0078] 1.1.1 A universal data identification system is used: The indicator requires that the data have a globally unique identifier, such as DOI, CSTR, Handle and other unique identifiers, and the identifier can be normally resolved and located to the corresponding data resource.

[0079] 1.1.2 Rich metadata is used to describe data: The indicator requires that metadata contain descriptive core elements such as creator, title, data identifier, publishing agency, publication date, abstract and keywords.

[0080] 1.1.3 Metadata accurately describes entity data: The indicator requirement is that the metadata describes the access address, file name, data volume and data format of the entity data.

[0081] 1.1.4 Metadata is provided in machine-readable format and can be machine-searched: The indicator requires whether the metadata is provided in machine-readable formats such as JSON-LD and RDFa.

[0082] 1.2 Three indicators and requirements are implemented in indicator A (accessibility):

[0083] 1.2.1 Metadata supports online user access: The indicator requirement is whether the metadata can be accessed through the HTTP protocol.

[0084] 1.2.2 Clear data openness and sharing methods: The indicator requirement is that the metadata includes the access level and access conditions of the data.

[0085] 1.2.3 Entity data supports user access: The indicator requirement is whether the entity data can be accessed through the HTTP / FTP protocol.

[0086] 1.3 Implemented one indicator item and indicator requirement in Indicator I (Interoperability):

[0087] 1.3.1 Use common interoperability protocols to provide data services: The indicator requirements are whether the data supports interface protocols such as RESTful, OAI-PMH, and the Chinese Academy of Sciences Data Center Interoperability Protocol to provide data services.

[0088] 1.4 Three indicators and indicator requirements are implemented in indicator R (reusability):

[0089] 1.4.1 Data usage permission: The indicator requirement is whether the metadata describes the data usage permission statement.

[0090] 1.4.2 Clear data ownership information supports data citation: The indicator requires that metadata describe the source information of data creation or generation.

[0091] 1.4.3 (Meta) Data complies with the standards of relevant scientific research fields: The indicator requirements are that the metadata uses common metadata standards and whether the entity data uses a data format that is common in the field.

[0092] 1.5 In indicator G (manageability), four indicators and indicator requirements were implemented:

[0093] 1.5.1 Data complies with security and ethical review requirements: The indicator requirements are that the data content complies with relevant legal provisions and does not contain inappropriate information that endangers public safety.

[0094] 1.5.2 Metadata complies with compliance requirements: Indicators are required to be described in core metadata specifications to support data discoverability.

[0095] 1.5.3 Use of independent identification system: The indicator requires the use of my country's independent identification system for data.

[0096] 1.5.4 Data traceability: The indicator requirement is that the metadata describes the entity information associated with the data.

[0097] 2. After building the FAIR evaluation index system, it is necessary to collect the scientific data to be evaluated, parse the data's unique identifiers, and generate standardized URLs. This information must be verified to ensure its validity. The collection process is as follows:

[0098] 2.1 For the DOI collection method, the following regular expression is used to determine validity and concatenate it into a DOI URL. This regular expression is used to match and validate DOIs. First, it matches DOIs that begin with doi:, http: / / doi.org / , or https: / / doi.org / . Next, it matches a DOI body that begins with 10. and is followed by a 4- to 9-digit digit, optionally followed by a version number. Finally, it matches all characters following a / to ensure the DOI's integrity.

[0099] (doi:\s*|(?:https?: / / )?doi.org / )? 10\.\d{4,9}(\.\d+)*( / .*)? $

[0100] 2.2 For the CSTR collection method, the following regular expression is used to determine whether it is valid and then concatenate it into a CSTRURL. This regular expression is used to match and verify CSTR. Its meaning is: First, it matches identifiers starting with CSTR: or BRID:, and supports English and Chinese colons or hyphens. Next, it matches identifiers consisting of five letters or numbers, followed by a period and a two-digit version number, and finally matches all subsequent characters.

[0101] ^((CSTR(:|:))|(BRID(:|:|-)))? [a-zA-Z0-9]{5}.\d{2}..*$

[0102] 2.3 Regarding the method of collecting URLs, determine whether it is valid by judging whether it can be parsed into a URL object.

[0103] 3. After completing the data collection in step 2 and obtaining the data standardized URL, the data metadata will be obtained and the metadata information will be aligned and verified. The detailed implementation method is as follows:

[0104] 3.1 Metadata Acquisition:

[0105] 3.1.1 Requesting Web Data: Use the Selenium tool to request the web address of the data identifier. To ensure that the information on the page is retrieved, set the parameter Wait Time (t) for JavaScript execution. If the time is too short, the data may not be fully retrieved, while if it is too long, the evaluation response speed will be affected. Based on testing, 10 seconds is the optimal setting for t in this implementation.

[0106] 3.1.2 Extract metadata: After the data in the page is loaded, parse the elements in the page DOM (Document Object Model), such as <meta> 、 <script type="application ld+json">,提取其中的可能存在元数据。

[0107] 3.1.3调用外部接口获取更多元数据:对于DOI,通过HTTP请求调用国际开放信息接口,例如Datacite API获取更多元数据;对于CSTR,通过HTTP请求调用国内自主信息接口,例如CSTR开放接口来获取更多的元数据;对于URL,则通过获取页面元数据中的DOI和CSTR再通过上述两个方式获取更多元数据。

[0108] 3.2多源元数据信息补齐与校验:

[0109] 3.2.1清洗:对获取到的元数据清理无关或冗余的字段,确保数据的精简性,例如对元数据中可能出现的特殊HTML标签进行去除,对特定元数据例如"keywords”进行类型检查。

[0110] 3.2.2元数据映射:根据预设的标准元数据词表对提取的元数据进行映射,将不同来源的元数据信息转换为统一键值对(key-value)的形式,便于后续统一处理。标准元数据词表通过Python代码构建,涵盖了各种标准字段名(如title、creator、summary等)。针对不同数据源的结构差异,例如DataCite元数据使用"title"表示数据名称,而CSTR元数据使用"name"表示数据名称,通过采用统一的映射,将这些异构元数据规范化为统一格式。通过这一过程,确保来自不同数据源的数据可以按照统一的标准进行管理和分析,从而消除由于元数据结构差异带来的不一致性。

[0111] 3.2.3合并:将不同来源的元数据合并为一个完整的元数据信息,如果在合并元数据的过程中发现某些key已经存在,则对其value的类型进行判断,如果值为字符串类型,则会根据相似度算法判断是否需要替换已有值;如果值为数组类型,则将新值与已有数组合并,并执行去重操作,确保数组中的元素的唯一性。经测试,在本实施中相似度算法使用token_sort_ratio为最优。

[0112] 3.2.4输出:经过以上步骤处理后的元数据将以标准化的键值对形式输出,为后续的FAIR评估提供数据支持。

[0113] 4.在完成步骤3得到的标准化的信息后,将对其进行FAIR能力的评估,按照步骤1构建的FAIR评估指标体系,分别对其中的五个指标进行计算评估,包含以下步骤:

[0114] 4.1加载评估规则:从数据库中读取步骤1构建的FAIR评估体系数据,用于后续计算评估。在本实施中,评估规则如下:

[0115] 4.1.1可发现

[0116] 4.1.1.1使用了通用的数据标识体系:检测数据标识符是否基于普遍接受的持久标识符方案,如DOI、Handle、CSTR等,检测数据标识符能够正常解析定位到数据集的主页面,本实施中该指标评分权重占比为40%。

[0117] 4.1.1.2采用了丰富的元数据描述数据:检测元数据对中是否全部包含核心元数据,本实施中该指标评分权重占比为30%,核心元数据包括数据创建者、标题、数据标识符、提交机构、发布机构、发布日期、摘要和关键字。

[0118] 4.1.1.3元数据准确描述了实体数据:检测元数据中包括实体数据访问地址、实体数据文件名称、数据量、数据格式字段,本实施中该指标评分权重占比为20%。

[0119] 4.1.1.4元数据提供机器可读格式并可被机器检索:检测元数据是否提供例如JSON-LD或RDFa等机器可读的格式,本实施中该指标评分权重占比为10%。

[0120] 4.1.2可访问

[0121] 4.1.2.1元数据支持用户在线访问:检测元数据是否可以通过HTTP方式访问,本实施中该指标评分权重占比为60%。

[0122] 4.1.2.2明确的数据开放共享方式:检测元数据中是否包含数据的访问级别和访问条件,如果访问级别是"有条件公开”,判断是否有限制条件;如果访问权限是"保护期”,判断是否有可用日期本。本实施中该指标评分权重占比为30%。

[0123] 4.1.2.3数据支持用户访问:若访问级别是完全公开,检测实体数据是否可以通过Http / Ftp方式访问,满足得1分,否则不得分。本实施中该指标评分权重占比为30%。

[0124] 4.1.3可互操作

[0125] 4.1.3.1采用通用的互操作协议提供数据服务:检测数据集发布方是否满足互操作协议。

[0126] 4.1.4可重用

[0127] 4.1.4.1使用数据许可:检测是否存在与数据的许可证信息相对应的元数据元素。本实施中该指标评分权重占比为50。

[0128] 4.1.4.2明确的数据权属信息支持数据引用:检测数据创建或生成的出处信息,例如作者,名称,版本(非必须),创建机构,创建时间,发布机构,发布时间,唯一标识符解析地址,本实施中该指标评分权重占比为30%。

[0129] 4.1.4.3(元)数据符合相关科研领域的标准:检测指标要求为元数据是否使用了通用元数据标准,例如中国科学院科学数据中心体系元数据规范,Dublin Core等。实体数据格式是否符合MIME类型(媒体类型),本实施中该指标评分权重占比为20%。

[0130] 4.1.5可管理

[0131] 4.1.5.1数据符合安全与伦理审查要求:检测数据中是否包含敏感词,本实施中该指标评分权重占比为40%。

[0132] 4.1.5.2元数据符合合规性要求:检测核心元数据元素中摘要字数是否满足100字以上,发布日期格式是否规范、关键字是否满足3个及以上、学科分类是否符合国内与国际标准,本实施中该指标评分权重占比为40%。

[0133] 4.1.5.3使用自主的标识体系:检测元数据中是否包含自主表示体系,满足则得1分,否则不得分,本实施中该指标评分权重占比为10%。

[0134] 4.1.5.4数据可溯源:检测元数据中是否包含项目关联信息,包括资助项目类型;项目编号;项目 / 课题名称;检测是否包含数据关联论文信息。本实施中该指标评分权重占比为10%。

[0135] 4.2评估分数:遍历五个维度的评估规则,计算五个维度中各指标项的分数,如下公式所示:

[0136]

[0137] 其中,X∈{F,A,I,R,G},i代表五个维度中各指标项的编号,Xtotal为每个维度规则中设定的满分,本实施中,Xtotal设定为100,即五个维度总分设定为100,为五个维度中的第i指标项的评分权重,为五个维度中第i指标项的评分。由此可得每个维度的评分,n为每个维度中指标项的数量,如下公式所示:

[0138]

[0139] 4.3计算平均分数:在得到每个维度的分数后,计算五个维度分数的平均数作为科学数据整体的FAIR能力完成度,计算公式如下:

[0140]

[0141] 其中S为科学数据整体的FAIR能力完成度。

[0142] 5.反馈

[0143] 5.1结果反馈:将步骤4中得到的计算评估结果中的分数建立可视化报表,便于后续分析。

[0144] 5.2提升建议:将步骤4中得到的计算评估结果中的各指标项与其总分对比,给出相关提升建议。提升建议包含以下几个方面:对于元数据完整性方面,针对评分较低的元数据项,尤其是缺失的关键元数据字段(如标题、创建者、描述等),建议补充完善相关内容。例如,添加数据集的标题、描述、创建者、联系方式、关键词等字段,以提高元数据的完整性。在数据互操作性方面,针对评分较低的互操作性指标,建议规范数据格式和结构。例如,建议采用标准化的数据格式(如CSV、JSON、XML等),确保数据在不同系统间的兼容性和可互操作性。在数据开放性方面,对于评分较低的数据开放性指标,建议增加合适的许可协议和访问权限设置,以便用户能够自由访问和使用数据。可以考虑使用开放许可证(如CC BY、CC0等)以及明确的访问权限设置,以促进数据的开放共享。在数据可用性方面,对于可用性评分较低的数据集,建议增强数据文档的可读性和易用性,提供详细的使用示例或者教程。

[0145] 本发明的另一实施例提供一种基于FAIR原则的科学数据评估系统,其包括:

[0146] 构建评估指标体系模块,用于基于FAIR原则构建科学数据的评估指标体系,所述评估指标体系包括可发现维度、可访问维度、可互操作维度、可重用维度和可管理维度,为每个维度定义若干评估指标,存储在数据库中;

[0147] 数据采集模块,用于采集待评估的科学数据,解析待评估的科学数据的唯一标识符,生成标准化的URL;

[0148] 数据处理模块,用于基于标准化的URL,获取数据的元数据信息并对其进行多源元数据信息对齐与校验;

[0149] 计算评估模块,用于基于对齐与校验的元数据信息和构建的评估指标体系,对每个维度的评估指标进行计算;

[0150] 反馈模块,用于基于评估指标计算结果,生成反馈报告。

[0151] 上述各模块的划分仅是举例说明,实际应用中可以根据需要而将上述功能分配由不同的功能模块完成,以完成前述方法中描述的全部或者部分功能。上述各模块的具体工作过程,可以参考前述方法实施例中的对应过程,在此不再赘述。

[0152] 本发明的另一实施例提供一种计算机设备(计算机、服务器、智能手机等),其包括存储器和处理器,所述存储器存储计算机程序,所述计算机程序被配置为由所述处理器执行,所述计算机程序包括用于执行本发明方法中各步骤的指令。

[0153] 本发明的另一实施例提供一种计算机可读存储介质(如ROM / RAM、磁盘、光盘),所述计算机可读存储介质存储计算机程序,所述计算机程序被计算机执行时,实现本发明方法的各个步骤。

[0154] 以上公开的本发明的具体实施例,其目的在于帮助理解本发明的内容并据以实施,本领域的普通技术人员可以理解,在不脱离本发明的精神和范围内,各种替换、变化和修改都是可能的。本发明不应局限于本说明书的实施例所公开的内容,本发明的保护范围以权利要求书界定的范围为准。< / script>

Claims

1. A scientific data evaluation method based on the FAIR principle, characterized by: The following steps are involved: Building an evaluation index system for scientific data based on the FAIR principle. The evaluation index system includes the dimensions of discoverability, accessibility, interoperability, reusability, and manageability. Several evaluation indicators are defined for each dimension and stored in a database. Collect scientific data to be evaluated, parse the unique identifiers of the scientific data to be evaluated, and generate standardized URLs; Based on standardized URLs, obtain metadata information of data and perform multi-source metadata alignment and verification; Based on the aligned and verified metadata information and the constructed evaluation indicator system, the evaluation indicators of each dimension are calculated; Generate a feedback report based on the evaluation indicator calculation results.

2. The scientific data evaluation method based on the FAIR principle according to claim 1 is characterized in that: The discoverable dimensions include the following evaluation metrics: Use a globally unique identifier that can be properly resolved and located on the data resource page; Metadata contains descriptive core elements including creator, title, data identifier, publishing agency, publication date, abstract, and keywords; Metadata accurately describes the access address, file name, data volume and data format of entity data; Metadata is provided in a machine-readable format and can be retrieved by machines.

3. The scientific data evaluation method based on the FAIR principle according to claim 1 is characterized in that: The accessibility dimension includes the following evaluation indicators: Metadata supports online access for users; The method of open data sharing is clear, including access levels and access conditions; Entity data supports user access.

4. The scientific data evaluation method based on the FAIR principle according to claim 1 is characterized in that: The interoperability dimension includes the following indicators: data supports common interoperability protocols.

5. The scientific data evaluation method based on the FAIR principle according to claim 1 is characterized in that: The reusability dimensions include the following evaluation indicators: Metadata describing the license statement for data use; Metadata clearly specifies the creation or generation source of data; The data meet the standards of relevant scientific research fields.

6. The scientific data evaluation method based on the FAIR principle according to claim 1 is characterized in that: The manageable dimensions include the following evaluation indicators: The data complies with safety and ethical review requirements; Metadata meets compliance requirements, and core metadata complies with domestic and international standards, such as abstract word count, keyword count, and subject classification; Use of self-labelling systems; The data supports traceability, and the metadata contains associated entity information and related items.

7. The scientific data evaluation method based on the FAIR principle according to claim 1 is characterized in that: The multi-source metadata information alignment and verification includes: Cleaning: Remove irrelevant or redundant fields from the acquired metadata to ensure data simplicity; Metadata mapping: Map the extracted metadata according to the preset standard metadata vocabulary, converting metadata information from different sources into a unified key-value pair format; Merge: Merge metadata from different sources into a complete metadata collection. If a key already exists during the metadata merging process, the value type is checked. If the value is a string, a similarity algorithm is used to determine whether the existing value needs to be replaced. If the value is an array, the new value is merged with the existing array and deduplication is performed to ensure the uniqueness of the elements in the array. Output: The metadata processed by the above steps will be output in the form of standardized key-value pairs.

8. The scientific data evaluation method based on the FAIR principle according to claim 1 is characterized in that: The metadata information based on alignment and verification and the constructed evaluation index system are used to calculate the evaluation index of each dimension, including: According to the constructed FAIR evaluation index system, the scores of each indicator in the five dimensions are calculated: Among them, X total The full score evaluation rules set in each dimension, X∈{F,A,I,R,G}, i is the specific indicator item number of each dimension, is the scoring weight of the i-th indicator item, is the score of the i-th indicator; thus, the score of each dimension is obtained: Where n is the number of specific indicators in each dimension; the average score of all dimensions is calculated according to the above formula: Where S is the overall FAIR capability completion of scientific data.

9. The scientific data evaluation method based on the FAIR principle according to claim 1, characterized in that: The feedback report includes evaluation scores, visualization reports and corresponding optimization suggestions.

10. A scientific data evaluation system based on the FAIR principle, characterized by: include: Construct an evaluation index system module, which is used to construct an evaluation index system for scientific data based on the FAIR principle. The evaluation index system includes the dimensions of discoverability, accessibility, interoperability, reusability, and manageability. Several evaluation indicators are defined for each dimension and stored in the database. The data collection module is used to collect scientific data to be evaluated, parse the unique identifier of the scientific data to be evaluated, and generate a standardized URL; The data processing module is used to obtain metadata information of the data based on the standardized URL and align and verify the multi-source metadata information; The calculation and evaluation module is used to calculate the evaluation indicators of each dimension based on the metadata information of alignment and verification and the constructed evaluation indicator system; The feedback module is used to generate a feedback report based on the evaluation indicator calculation results.