A Method and System for Identifying and Tracing Scientific and Technological Data throughout the Lifecycle

By combining the evidence storage server and identification server with blockchain and large language model, the full life cycle identification and traceability of scientific and technological data is solved, the authenticity and traceability of scientific research data are realized, and automatic identification and description functions are provided.

CN119848636BActive Publication Date: 2025-08-05ZHEJIANG TOPCHEER INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510336616.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-08-05
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

The existing technology lacks effective methods for identifying and traceability of scientific and technological data throughout the life cycle, making it difficult to ensure the authenticity and traceability of scientific and technological data, especially in scientific research data, there is a risk of data tampering.

Method used

By setting up an evidence storage server and an identification server, using data fingerprints and data identification to generate evidence storage information, and using blockchain evidence storage, combining large language models to identify data to achieve automatic identification and traceability of scientific research data.

Benefits of technology

It provides guarantees for the authenticity of scientific research data throughout the life cycle, can automatically identify data types and give data descriptions, ensure the authenticity and traceability of data, and is suitable for the evidence-keeping and identification of scientific research data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119848636B_ABST
    Figure CN119848636B_ABST
Patent Text Reader

Abstract

The multiple embodiments of this specification relate to, in particular, a method and system for identifying and tracing scientific and technological data throughout its life cycle. A method for identifying and tracing scientific and technological data throughout its life cycle comprises the following steps: setting up an evidence server and an identification server; the evidence server receives the data fingerprint, data identifier, and data description of scientific research data; the evidence server generates evidence information for each scientific research data; packaging multiple pieces of evidence information into an evidence package at a preset period, extracting the data fingerprint of the evidence package and storing it; the identification server receives scientific research data and data descriptions of the scientific research data, uses the data descriptions as labels for the scientific research data, and generates sample data; inputs the sample data into a pre-connected large language model, and obtains a data recognition model based on the large language model; obtains a data description for identifying the scientific research data based on the response of the data recognition model to the scientific research data to be identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Multiple embodiments of this specification relate to, specifically, a method and system for identifying and tracing scientific and technological data throughout its life cycle. Background Art

[0002] Scientific and technological data refers to various types of information and materials generated during the process of scientific research and technological development. This information can be data records obtained through methods such as experiments, observations, surveys, and simulations, or it can be content obtained through preliminary data processing and analysis. It can also be data directly output by instruments and equipment. Scientific and technological data includes not only numerical data such as measurements and statistical results, but also content in various forms such as text, images, audio, and video. Scientific and technological data is the foundation of scientific research and an important basis for the approval and completion of relevant scientific and technological projects. Therefore, ensuring the authenticity of scientific and technological data and achieving traceability of this data is of paramount importance. Given the wide variety of scientific and technological data, manual identification of its types is difficult. Currently, there is a lack of effective technology for tracing and identifying the source of scientific and technological data throughout its life cycle. Summary of the Invention

[0003] Multiple embodiments of this specification describe a method and system for identifying and tracing scientific and technological data throughout its life cycle.

[0004] In a first aspect, the embodiments of this specification provide a method for identifying and tracing scientific and technological data throughout its life cycle, including the following steps:

[0005] Set up the evidence storage server and identification server;

[0006] The evidence storage server receives data fingerprints, data identifiers, and data descriptions of scientific research data, wherein the scientific research data includes output data of scientific research equipment and research data;

[0007] The evidence storage server generates evidence information for each scientific research data, and the evidence information includes the data fingerprint and data identifier;

[0008] Packing the plurality of evidence information into an evidence package at a preset period, extracting the data fingerprint of the evidence package and storing it through the blockchain, and storing the evidence package;

[0009] The recognition server receives scientific research data and a data description of the scientific research data, uses the data description as a label of the scientific research data, and generates sample data;

[0010] Inputting the sample data into a pre-connected large language model, and obtaining a data recognition model based on the large language model;

[0011] According to the response of the data identification model to the scientific research data to be identified, a data description for identifying the scientific research data is obtained.

[0012] In a second aspect, the embodiments of this specification provide a full life cycle technology data identification and traceability system, including:

[0013] Evidence server and identification server,

[0014] The evidence server receives data fingerprints, data identifiers, and data descriptions of scientific research data, the scientific research data including output data of scientific research equipment and research data, generates evidence information for each scientific research data, the evidence information including the data fingerprint and data identifier, packages multiple pieces of evidence information into evidence packages at a preset period, extracts the data fingerprints of the evidence packages, stores them through blockchain, and stores the evidence packages;

[0015] The recognition server receives scientific research data and a data description of the scientific research data, uses the data description as a label for the scientific research data, generates sample data, inputs the sample data into a pre-connected large language model, obtains a data recognition model based on the large language model, and obtains a data description for identifying the scientific research data based on the response of the data recognition model to the scientific research data to be identified.

[0016] In a third aspect, embodiments of this specification provide an electronic device, including a processor and a memory;

[0017] The processor is connected to the memory;

[0018] The memory is used to store executable program code;

[0019] The processor reads the executable program code stored in the memory to run a program corresponding to the executable program code, so as to execute the method described in any one of the above aspects.

[0020] In a fourth aspect, an embodiment of this specification provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method described in any one of the above aspects is implemented.

[0021] In a fifth aspect, embodiments of this specification provide a computer program product, including a computer program, which implements the method described in any of the above aspects when executed by a processor.

[0022] The beneficial effects of the technical solutions provided by some embodiments of this specification include at least:

[0023] In various embodiments of this specification, a full-lifecycle scientific and technological data identification and traceability method and system are provided. By configuring a data concentrator to collect data output by scientific research equipment and manually input through scientific research terminals, this system then implements the storage, traceability, and identification of this scientific research data through a storage server and an identification server. This system ensures the authenticity of the scientific research data and automatically identifies the type of given scientific research data and provides a data description.

[0024] Other features and advantages of the various embodiments of this specification will be further disclosed in the following detailed description and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of this specification, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0026] Figure 1 Schematic diagram of the application scenario of the data identification and traceability method provided in the embodiments of this specification.

[0027] Figure 2 A flow chart of the data identification and traceability method provided in the embodiments of this specification.

[0028] Figure 3 Schematic diagram of data storage provided in the embodiments of this specification.

[0029] Figure 4 Flowchart of the method for receiving scientific research data by the evidence storage server.

[0030] Figure 5 Flowchart of the method for obtaining data identification model.

[0031] Figure 6 This is a schematic diagram of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION

[0032] The following is an explanation and description of the technical solutions of the embodiments of this specification in conjunction with the drawings of the embodiments of this specification. However, the following embodiments are only preferred embodiments of this specification and are not exhaustive. Based on the embodiments in the implementation mode, other embodiments obtained by those skilled in the art without making any creative work are all within the scope of protection of this specification.

[0033] Throughout this specification, the claims, and the accompanying drawings, the terms "first," "second," "third," and the like are used to distinguish between different items, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may include other steps or elements inherent to the process, method, product, or apparatus.

[0034] In the following description, terms such as "inside", "outside", "up", "down", "left", "right", etc. that indicate directions or positional relationships are only used to facilitate the description of the embodiments and simplify the description. They do not indicate or imply that the devices or elements referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, they should not be understood as limitations on this specification.

[0035] The data involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection of relevant data complies with the relevant laws, regulations and standards of relevant countries and regions.

[0036] Before introducing the technical solution recorded in this specification, the application scenarios of the technical solution and related technologies are introduced.

[0037] The authenticity of scientific research data is crucial during project review. It is the cornerstone of scientific research and technological development, directly impacting the reliability of research conclusions and the validity of scientific discoveries. Distortion of scientific research data can not only lead to erroneous conclusions but also mislead subsequent research directions, waste valuable research resources, and even impact public policymaking and social development decisions. Therefore, ensuring data authenticity is a core element in assessing project quality during project review.

[0038] When reviewers examine projects, they use the submitted data to determine whether the research methods are scientifically sound, the experimental results are convincing, and the theoretical hypotheses are valid. Only with authentic data can accurate research conclusions be drawn. Researchers must be responsible for their work, honestly record and report experimental results, and avoid any form of data fabrication or tampering. Data fabrication incidents can not only harm the careers of the individuals involved, but also negatively impact the reputation of the entire research institution.

[0039] Ensuring data authenticity is particularly important in research involving sensitive areas such as health and the environment. For example, in drug development, the authenticity of clinical trial data directly impacts the safety and efficacy evaluation of new drugs, ultimately determining their marketability and patient benefit. Therefore, strictly controlling data authenticity during the project review phase is a socially responsible approach and a key step in safeguarding the public interest.

[0040] Scientific research data 21 includes the output data of scientific research equipment 13, that is, raw data, which refers to the unprocessed initial observations or measurements. The output data of scientific research equipment 13 comes directly from experimental, observational and other equipment. This type of data retains the most original information. Data after preliminary processing refers to the results obtained after cleaning, converting, summarizing and other operations on the original data. Text data covers various written records, such as research reports and papers. Data in the form of images, audio, video, etc. are also becoming more and more common in modern scientific research. This type of unstructured data often contains rich visual or auditory information and can be used in multiple fields such as biometrics, speech processing, and natural scene understanding. These data are directly processed and obtained by scientific researchers or project personnel and then input into the receiving device.

[0041] Scientific research data 21 not only needs to be authentic and traceable, but also needs to be studied due to the wide variety of scientific research data 21. To this end, this application proposes a full life cycle scientific and technological data identification and traceability method and system, please refer to the attached Figure 1 This system utilizes a data concentrator 11 to collect data output from scientific research equipment 13 and input manually through scientific research terminals 12. The system then uses a storage server 20 and an identification server 30 to perform archival, traceability, and identification of scientific research data 21. This system ensures the authenticity of scientific research data 21. Furthermore, for a given piece of scientific research data 21, the system automatically identifies its type and provides a data description. For example, the data description indicates that it is generated by automatic monitoring equipment at a weather station and is commonly used in climate change research.

[0042] An interactive interface is provided on the scientific research terminal 12. The interactive interface is used to upload research data. Or to request identification or tracing of given scientific research data 21. Among them, the interactive interface can be run on the terminal device in the form of a browser, or can be run on the terminal device in the form of an independent application, etc. The specific presentation form of the client is not limited here. The server involved in this application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms. The scientific research terminal 12 can be a smart phone, tablet computer, laptop computer, PDA, personal computer, smart speaker, smart TV, smart watch, car-mounted equipment, wearable device, etc., but is not limited to this. The scientific research terminal 12 and the server can be directly or indirectly connected by wired or wireless communication, and this application does not limit it here. The number of servers and scientific research terminals 12 is also not limited.

[0043] This manual first provides a method for identifying and tracing scientific and technological data throughout its life cycle. Figure 2 , including the steps of:

[0044] Step S101) Set up the evidence storage server 20 and the identification server 30. Setting up the evidence storage server 20 and the identification server 30 to perform different tasks separately achieves business separation, allowing the system to carry a larger load, with minimal mutual impact between businesses, and higher security.

[0045] In step S102, the evidence storage server 20 receives the data fingerprint 24, data identifier 25, and data description of the scientific research data 21. The scientific research data 21 includes the output data of the scientific research equipment 13 and the research data. The data fingerprint 24 enables the scientific research data 21 to be stored as evidence, providing proof of authenticity for the scientific research data 21. A data fingerprint 24 (sometimes also called a data hash or content fingerprint) is a fixed-length string generated by processing a piece of data using a specific algorithm. This string is a unique identifier for the original data and can be used to quickly identify and verify data integrity or detect duplicate data. Data fingerprints 24 are widely used in various fields, including file systems, database management, network security, and software development. Data fingerprints 24 can also serve as index keys to accelerate the search process, especially in distributed file systems. Methods for generating data fingerprints 24 typically rely on hash functions, such as MD5, SHA-1, and SHA-256. These functions accept data input of arbitrary length and output a fixed-length hash value.

[0046] Step S103 ) The evidence server 20 generates evidence information for each scientific research data 21 , and the evidence information includes the data fingerprint 24 and the data identifier 25 .

[0047] The scientific research data 21 can be stored as evidence through the data fingerprint 24 and data identifier 25. The method for verifying the data fingerprint 24 is to use a predetermined extraction function to extract the data fingerprint 24 of the scientific research data 21 and compare it with the data fingerprint 24 stored on the blockchain. If they match, it indicates that the scientific research data 21 has not been tampered with and is authentic. Conversely, if they do not match, it indicates that the scientific research data 21 has been tampered with.

[0048] In another improved embodiment, the evidence server 20 and the identification server 30 continuously exchange evidence codes 22 in pairs. The evidence information includes the data fingerprint 24, the data identifier 25, a pair of evidence codes 22, and a supplementary code 23. The method by which the evidence server 20 generates the supplementary code 23 includes: continuously attempting to generate the supplementary code 23 until the supplementary code 23 satisfies the numerical continuity of the last preset digits of the hash value of the evidence information of the scientific research data 21. For reasons of space, the last preset digits are assumed to be the last three digits, and the hash value is expressed in hexadecimal. Assume that the hash value of the first scientific research data 21 is: 846E8C38ADF75840D6FB72E2B35A624CB1625D49F6CFB8D65A931D0029884354, the pair of evidence codes 22 are: 4ED12B, 64C87A, and the supplementary code 23 is: 82D5AC. Then, after splicing the hash value of the scientific research data 21, the pair of evidence codes 22, and the supplementary code 23, we can extract the “ The hash value of "846E8C38ADF75840D6FB72E2B35A624CB1625D49F6CFB8D65A931D00298843544ED12B64C87A82D5AC" is 92CDB42B62F6A62034EBB14E83BBC987F00B6E7D02596B38274C1C5D7D5254C0. That is, the last three digits are "4C0", which means that the last three digits of the hash value of the next scientific research data 21, a pair of evidence codes 22, and the supplementary code 23 must be "4C1". Since the hash function is almost random, the probability of obtaining a hash value with the last three digits of "4C1" is 1 / (16^3), or 1 / 4096. This means that the supplementary code 23 is randomly generated, and there is a 1 / 4096 probability that the last three digits of the hash value will be consecutive. This means that the random generation of the supplementary code 23 is performed 4096 times for each scientific research data 21. The difficulty increases when more digits are required at the end.

[0049] Since blockchain-based evidence storage requires initiating a transfer transaction on the blockchain, the information to be stored must be included in the transaction notes. Transactions incur a handling fee, and frequent storage consumes significant funds. To save this cost, the frequency of storage must be reduced. However, in the time between storages, scientific research data 21 can be tampered with without detection. Therefore, the frequency of storage cannot be too low. To this end, this embodiment utilizes the requirements of the supplementary code 23 to implement special processing between two storages. This ensures that any tampering with scientific research data 21 requires the immediate generation of a new supplementary code 23. If the required number of consecutive trailing digits is sufficient, theoretically, the time required to attempt to generate a supplementary code 23 should exceed the time interval between two storages, thus ensuring the authenticity of each scientific research data 21. The total number of digits in the pair of the evidence code 22 and the supplementary code 23 is a preset length. The supplementary codes 23 are stored continuously. Since the pair of evidence codes 22 is stored separately on the evidence server 20 and the identification server 30, tampering with scientific research data 21 requires modifications not only to the evidence server 20 but also to the identification server 30. When the management or modification permissions of two servers are controlled by different entities, it is more difficult to tamper with scientific research data21.

[0050] Step S104) The multiple pieces of evidence information are packaged into an evidence package at a preset period. The data fingerprint 24 of the evidence package is extracted and recorded via the blockchain, and the evidence package is stored. When using a supplementary code 23 and requiring the trailing preset bits of the hash value of the evidence information of the scientific research data 21 to be numerically continuous, the period can be increased, reducing the number of blockchain recordings, while still ensuring the same data authenticity.

[0051] On the other hand, please see the attached Figure 4 The method for the evidence storage server 20 to receive the data fingerprint 24, data identifier 25 and data description of the scientific research data 21 includes:

[0052] Step S201) Setting a data concentrator 11 for a plurality of scientific research equipment 13, wherein the data concentrator 11 receives output data of the scientific research equipment 13;

[0053] Step S202) The data concentrator 11 packages the received output data of the scientific research equipment 13 into scientific research data 21 at a preset time period, and extracts the data fingerprint 24 of the scientific research data 21;

[0054] Step S203) Generate a data identifier 25 for the scientific research data 21 according to a preset identifier generation rule;

[0055] Step S204) generating a data description based on the manually preset description of the scientific research equipment 13 and the current status setting of the scientific research equipment 13;

[0056] Step S205 ) The data concentrator 11 receives the manually submitted research data and data description, extracts the data fingerprint 24 of the research data, and generates a data identifier 25 for the research data;

[0057] Step S206) The data concentrator 11 periodically reports the data fingerprint 24, data identifier 25, and data description of the scientific research data 21 to the evidence storage server 20. By configuring the data concentrator 11 to collect the output data of the scientific research equipment 13 and perform evidence storage operations on the data concentrator 11, the security of the data concentrator 11 is controlled, thereby achieving higher security.

[0058] As a recommended approach, the data concentrator 11 is a trusted device. Trusted devices refer to hardware and software systems that are recognized as safe, reliable and capable of ensuring data integrity and privacy protection. These devices play a vital role in the entire life cycle of scientific research data 21. From data collection, transmission, storage to final analysis and sharing, they ensure data quality and security at every link. The design and implementation of trusted devices rely on a series of technical and management measures. First, at the hardware level, such devices are usually equipped with dedicated security chips or modules for encryption and decryption, identity authentication, and integrity protection of key operations.

[0059] Step S105 ) The recognition server 30 receives the scientific research data 21 and the data description of the scientific research data 21 , uses the data description as a label of the scientific research data 21 , and generates sample data.

[0060] Step S106) Input the sample data into a pre-connected large language model, and obtain a data recognition model based on the large language model. The method of inputting the sample data into the pre-connected large language model includes:

[0061] Constructing a natural language task according to the sample data to instruct the large language model to learn the sample data;

[0062] Submitting the natural language task to the large language model.

[0063] Step S107 ) According to the response of the data recognition model to the scientific research data 21 to be identified, a data description for identifying the scientific research data 21 is obtained.

[0064] Please see the attached Figure 5 The method for obtaining a data recognition model based on the large language model includes:

[0065] Step S301) Construct a natural language task to generate extended samples based on the sample data and submit it to the large language model to obtain extended samples. First, natural language tasks are created based on the existing sample data. These tasks aim to generate more samples, known as extended samples. These tasks can be designed for tasks such as blank filling, question-answer pair generation, and text summarization. The goal is to leverage the power of the large language model to create new samples that are similar to but distinct from the original samples. After submitting these natural language tasks to the pre-trained large language model, the model generates a series of new samples based on the task requirements, thereby enriching the training set.

[0066] Step S302) A neural network model is established as a student model, and the large language model is used as a teacher model. The student model is typically a simpler model with fewer parameters. Its goal is to learn by imitating the behavior of the teacher model. This may reduce accuracy, but it is more efficient and consumes fewer resources.

[0067] Step S303) The student model learns the teacher model on the sample data and the extended sample. This can be achieved in a variety of ways, such as directly copying the teacher model's prediction results as labels, or having the student model attempt to imitate the activation state of the teacher model's intermediate layers.

[0068] Step S304) A data recognition model is obtained based on the learned neural network model. This model maintains low computational cost while approaching the performance of the teacher model. When used in real-world applications, it can quickly respond to user queries and perform large-scale data analysis without impacting user experience or business efficiency due to resource constraints.

[0069] On the other hand, this manual provides a full life cycle scientific and technological data identification and traceability system, including:

[0070] Evidence server 20 and identification server 30,

[0071] The evidence server 20 receives a data fingerprint 24, a data identifier 25, and a data description of scientific research data 21, wherein the scientific research data 21 includes output data of the scientific research equipment 13 and research data, generates evidence information for each scientific research data 21, wherein the evidence information includes the data fingerprint 24 and the data identifier 25, packages multiple pieces of evidence information into an evidence package at a preset period, extracts the data fingerprint 24 of the evidence package, stores the evidence through the blockchain, and stores the evidence package;

[0072] The recognition server 30 receives the scientific research data 21 and the data description of the scientific research data 21, uses the data description as a label of the scientific research data 21, generates sample data, inputs the sample data into a pre-connected large language model, obtains a data recognition model based on the large language model, and obtains a data description for identifying the scientific research data 21 based on the response of the data recognition model to the scientific research data 21 to be identified.

[0073] See also Figure 6 A schematic structural diagram of an electronic device provided in an embodiment of this specification is shown.

[0074] like Figure 6 As shown, the electronic device 1100 may include: at least one processor 1101, at least one network interface 1104, a user interface 1103, a memory 1105, and at least one communication bus 1102. The communication bus 1102 may be used to implement communication between the aforementioned components. The user interface 1103 may include buttons, and optionally may also include a standard wired interface or a wireless interface. The network interface 1104 may include, but is not limited to, a Bluetooth module, an NFC module, a Wi-Fi module, etc. The processor 1101 may include one or more processing cores. The processor 1101 utilizes various interfaces and circuits to connect the various components within the entire electronic device 1100. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 1105 and accessing data stored in the memory 1105, it performs various functions of the routing device 1100 and processes data. Optionally, the processor 1101 may be implemented in hardware using at least one of a DSP, an FPGA, and a PLA. The processor 1101 may integrate one or a combination of a CPU, a GPU, and a modem. The CPU primarily processes the operating system, user interface, and applications; the GPU is responsible for rendering and drawing content displayed on the display; and the modem handles wireless communications.

[0075] It is understandable that the above-mentioned modem may not be integrated into the processor 1101, but may be implemented separately through a chip.

[0076] Memory 1105 may include either RAM or ROM. Optionally, memory 1105 may include non-transitory computer-readable media. Memory 1105 may be used to store instructions, programs, codes, code sets, or instruction sets. Memory 1105 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, sound playback function, image playback function, etc.), instructions for implementing the aforementioned method embodiments, etc.; the data storage area may store data related to the aforementioned method embodiments, etc. Memory 1105 may also optionally be at least one storage device located remotely from the aforementioned processor 1101. Memory 1105, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and application programs. Processor 1101 may be configured to invoke the application programs stored in memory 1105 and execute the methods described in the aforementioned embodiments.

[0077] The embodiments of this specification also provide a computer-readable storage medium having instructions stored therein that, when executed on a computer or processor, cause the computer or processor to perform the steps of the aforementioned embodiments. If the components of the aforementioned electronic device are implemented as software functional units and sold or used as independent products, they may be stored in the computer-readable storage medium.

[0078] The embodiments of this specification also provide a computer program product, including a computer program, which implements multiple steps in the above embodiments when executed by a processor.

[0079] In the absence of conflict, the technical features in this embodiment and implementation scheme can be combined arbitrarily.

[0080] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product comprises multiple computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this specification are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium accessible by a computer or a data storage device such as a server or data center that integrates multiple available media. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a digital versatile disc (DVD)), or a semiconductor medium (eg, a solid state drive (SSD)).

[0081] When implemented via hardware or firmware, the aforementioned method flow is programmed into the hardware circuit to obtain the corresponding hardware circuit structure and realize the corresponding function. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit, whose logical function is determined by the user's device programming. Designers can "integrate" a digital system on a PLD through self-programming, eliminating the need for chip manufacturers to design and manufacture dedicated integrated circuit chips. Moreover, today, instead of manually manufacturing integrated circuit chips, this programming is often performed using "logic compiler" software. This is similar to the software compiler used in program development. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There are not just one HDL, but many. Those skilled in the art will also understand that simply by programming the method flow in one of the aforementioned hardware description languages and programming it into the integrated circuit, a hardware circuit that implements the logical method flow can be easily obtained.

[0082] The embodiments described above are merely preferred embodiments of this specification and are not intended to limit the scope of this specification. Without departing from the design spirit of this specification, various modifications and improvements made to the technical solutions of this specification by ordinary technicians in this field should fall within the scope of protection determined by the claims of this specification.

Claims

1. A method for identifying and tracing scientific and technological data throughout its life cycle, characterized in that: Including steps: Set up the evidence storage server and identification server; The evidence storage server receives data fingerprints, data identifiers, and data descriptions of scientific research data, wherein the scientific research data includes output data of scientific research equipment and research data; The evidence storage server generates evidence information for each scientific research data, and the evidence information includes the data fingerprint and data identifier; Packing the plurality of evidence information into an evidence package at a preset period, extracting the data fingerprint of the evidence package and storing the evidence through the blockchain, and storing the evidence package; The recognition server receives scientific research data and a data description of the scientific research data, uses the data description as a label of the scientific research data, and generates sample data; Inputting the sample data into a pre-connected large language model, and obtaining a data recognition model based on the large language model; Obtaining a data description for identifying the scientific research data according to a response of the data identification model to the scientific research data to be identified; The evidence server and the identification server continuously exchange evidence codes in pairs. The evidence information includes the data fingerprint, data identifier, a pair of evidence codes and a supplementary code. The method for generating the supplementary code by the evidence storage server includes: The supplementary code is continuously attempted to be generated until the supplementary code satisfies the preset position at the end of the hash value of the evidence information of the scientific research data, and is numerically continuous.

2. A full life cycle scientific and technological data identification and traceability method according to claim 1, characterized in that: The method for the evidence storage server to receive the data fingerprint, data identifier and data description of scientific research data includes: Setting a data concentrator for a plurality of scientific research equipment, wherein the data concentrator receives output data of the scientific research equipment; The data concentrator packages the received output data of the scientific research equipment into scientific research data at a preset time period and extracts the data fingerprint of the scientific research data; Generate a data identifier for the scientific research data according to a preset identifier generation rule; Generate data description based on the manually preset scientific research equipment description and the current scientific research equipment status settings; The data concentrator receives manually submitted research data and data descriptions, extracts data fingerprints of the research data, and generates data identifiers for the research data; The data concentrator periodically reports the data fingerprint, data identification and data description of the scientific research data to the evidence storage server.

3. A full life cycle scientific and technological data identification and traceability method according to claim 2, characterized in that: The data concentrator is a trusted device.

4. A method for identifying and tracing scientific and technological data throughout its life cycle according to claim 1 or 2, characterized in that: The method of inputting the sample data into the pre-connected large language model includes: Constructing a natural language task according to the sample data to instruct the large language model to learn the sample data; Submitting the natural language task to the large language model.

5. A method for identifying and tracing scientific and technological data throughout its life cycle according to claim 1 or 2, characterized in that: The method for obtaining a data recognition model according to the large language model includes: Constructing a natural language task for generating extended samples based on the sample data, and submitting the task to the large language model to obtain extended samples; Establishing a neural network model as a student model and using the large language model as a teacher model; On the sample data and the extended samples, enabling the student model to learn the teacher model; A data recognition model is obtained based on the learned neural network model.

6. A full life cycle scientific and technological data identification and traceability system, characterized by: include: Evidence server and identification server, The evidence server receives data fingerprints, data identifiers, and data descriptions of scientific research data, the scientific research data including output data of scientific research equipment and research data, generates evidence information for each scientific research data, the evidence information including the data fingerprint and data identifier, packages multiple pieces of evidence information into evidence packages at a preset period, extracts the data fingerprints of the evidence packages, stores them through blockchain, and stores the evidence packages; The recognition server receives scientific research data and a data description of the scientific research data, uses the data description as a label for the scientific research data, generates sample data, inputs the sample data into a pre-connected large language model, obtains a data recognition model based on the large language model, and obtains a data description for identifying the scientific research data based on a response of the data recognition model to the scientific research data to be identified; The evidence server and the identification server continuously exchange evidence codes in pairs. The evidence information includes the data fingerprint, data identifier, a pair of evidence codes and a supplementary code. The method for generating the supplementary code by the evidence storage server includes: The supplementary code is continuously attempted to be generated until the supplementary code satisfies the preset position at the end of the hash value of the evidence information of the scientific research data, and is numerically continuous.

7. An electronic device, characterized in that including a processor and a memory; The processor is connected to the memory; The memory is used to store executable program code; The processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to execute the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

9. Computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Block chain-based scientific research data evidence storage method, computer system and storage medium

    CN116341021A

  • Article information analysis model generation method and device and electronic equipment

    CN118864042A