Archive retrieval method and system based on block chain

By splitting electronic files into index information and full-text information on the blockchain, and using smart contracts to calculate the degree of matching, the problem of unrelated archive traversal in the existing blockchain search methods is solved, and efficient and accurate archive retrieval and data security are achieved.

CN120011472APending Publication Date: 2025-05-16CHINA SOUTHERN POWER GRID CO LTD SHARED OPERATION CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411605415.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-16
Filing Date
2024-11-12
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

After the existing blockchain search method enters keywords, it needs to read all the archive information stored on the chain, resulting in traversing unrelated archives, which causes great inconvenience and is not refined enough.

Method used

A blockchain-based archive search method is proposed. By splitting electronic archives into index information and full-text information, it is stored on the index chain and full-text information chain respectively, and the smart contract is used to calculate the degree of matching the search request with the archive, and the most relevant archive index information and full-text information are returned.

Benefits of technology

Improve the search efficiency and accuracy, and users can quickly obtain the list of archives most relevant to the search request, reduce the traversal of unrelated archives, enhance the security and transparency of data, and improve the overall user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011472A_ABST
    Figure CN120011472A_ABST
Patent Text Reader

Abstract

The invention provides an archive retrieval method and system based on a block chain. The method comprises the following steps: forming an electronic file of an archive, namely an electronic archive, splitting the electronic archive into index information and full-text information, and respectively placing the index information and the full-text information on an index chain and a full-text information chain; receiving a user retrieval request, calculating the matching degree of the retrieval request and the archive entry through the smart contract, and returning a list of the index information of the archive most related to the user retrieval request; and positioning full-text information of the electronic archive according to the index information, verifying the archive through the smart contract, and returning an archive full-text information list which is verified successfully and is most related to the matching degree of the retrieval condition to the user after the verification is passed. By means of the method and the corresponding system, the retrieval speed can be increased, and more accurate retrieval results can be provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information processing, and in particular to a blockchain-based archive retrieval method and system. Background Art

[0002] Blockchain is a new application model of distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm computer technology. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block.

[0003] At present, in the process of searching the blockchain, after the user enters a keyword, it is usually necessary to read all the information of the files stored on the chain, and some irrelevant files also need to be traversed in their entirety, which brings great inconvenience, or the current search is not refined enough. Summary of the invention

[0004] The present invention provides a blockchain-based archive retrieval method and system to solve the above-mentioned problems:

[0005] The present invention proposes a blockchain-based archive retrieval method, the method comprising:

[0006] The electronic files forming the archives are electronic archives, which are divided into index information and full-text information, and are placed on the index chain and the full-text information chain respectively;

[0007] Receive a user search request, calculate the matching degree between the search request and the archive entry through a smart contract, and return a list of index information of the archives most relevant to the user search request;

[0008] The full-text information of the electronic archive is located according to the index information, and the archive is verified through the smart contract. Once the verification is successful, a list of full-text archive information that is successfully verified and most relevant to the search criteria is returned to the user.

[0009] Furthermore, a blockchain-based archive retrieval method forms an electronic file of the archive, namely an electronic archive, splits the electronic archive into index information and full-text information, and places them on the index chain and the full-text information chain respectively, including:

[0010] Forming electronic archives, extracting index information from each electronic archive, the index information including: the number of times the archive has been queried, the length of the electronic archive, the archive timestamp, the frequency of modification of the archive, the number of times the archive has been cited, the number of times the archive has been queried, the time of the most recent modification of the archive and the archive summary, and placing the index information on an index chain;

[0011] Calculate the hash value of the electronic file, place the full-text information of the electronic file together with the hash value on the information chain, and retain a pointer to the corresponding full-text information in the information chain in each block of the index chain.

[0012] Furthermore, a blockchain-based archive retrieval method receives a user retrieval request, calculates the matching degree between the retrieval request and the archive entry through a smart contract, and returns a list of archives most relevant to the user retrieval request, including:

[0013] Receiving a search request from a user, wherein the search request includes a keyword or phrase that the user wishes to search for;

[0014] Preprocessing the search request, wherein the preprocessing includes removing stop words and segmenting words;

[0015] The smart contract calculates the matching degree between the index information on the index chain and the user's search request through the archive relevance model;

[0016] According to the matching degree calculation results, the index information entries of all candidate archives are sorted through the index optimization model to ensure that the index information entries of the most relevant archives are ranked first.

[0017] Furthermore, in a blockchain-based archive retrieval method, the archive relevance model includes:

[0018] The archive correlation model is:

[0019]

[0020] Among them, W is the relevance of the archive, f is the frequency of keywords in the abstract of the electronic archive index information, t0 is the time when the retrieval request is received, t is the timestamp of the archive, L is the length of the electronic archive, and α, β and θ are weight coefficients.

[0021] Furthermore, in a blockchain-based archive retrieval method, the index optimization model includes:

[0022] The index optimization model is:

[0023]

[0024] Among them, P represents the sorting priority, λ, μ and ν are the weight coefficients that control the contribution of each part and can be adjusted according to the retrieval requirements. q and w s They represent the number of file queries and the weight of user preference, q and s represent the number of queries and user preference statistics, Q max and S maxare the maximum expected values ​​of q and s, respectively. T represents the freshness of the document, that is, the time from the last modification to now. T max is the maximum value of T and is used to normalize T. MR is a combined factor that takes into account the modification frequency and the number of citations. max is its maximum expected value, and κ is a nonlinear coefficient used to adjust the influence of MR.

[0025] Furthermore, a blockchain-based archive retrieval method locates the full-text information of the electronic archive according to the index information, verifies the archive through the smart contract, and returns a list of the full-text information of the archive that is successfully verified and most relevant to the search condition to the user, including:

[0026] Locate the full-text information of the electronic archive according to the pointer to the full-text information in the index information;

[0027] Compare the hash value of the archive content with the hash value stored on the blockchain through the smart contract;

[0028] If the comparison is consistent, the full text information list of the archive is returned to the user.

[0029] The present invention proposes a blockchain-based archive retrieval system, the system comprising:

[0030] Generate an electronic archive module, which is used to form an electronic file of the archive, namely, an electronic archive, split the electronic archive into index information and full-text information, and place them on the index chain and the full-text information chain respectively;

[0031] A retrieval index information module is used to receive a user retrieval request, calculate the matching degree between the retrieval request and the archive entry through a smart contract, and return a list of index information of archives most relevant to the user's retrieval request;

[0032] The full-text information return module is used to locate the full-text information of the electronic archive according to the index information, and verify the archive through the smart contract. If the verification is successful, a list of full-text information of the archives that are successfully verified and most relevant to the search criteria is returned to the user.

[0033] Furthermore, in a blockchain-based archive retrieval system, the electronic archive generation module includes:

[0034] Generate an index chain module, which is used to form an electronic archive, extract index information from each electronic archive, the index information includes: the number of times the archive is queried, the length of the electronic archive, the archive timestamp, the modification frequency of the archive, the number of times the archive is cited, the number of times the archive is queried, the time of the most recent modification of the archive and the archive summary, and place the index information on the index chain;

[0035] Generate an information chain module, which is used to calculate the hash value of the electronic file, place the full-text information of the electronic file and the hash value together on the information chain, and retain a pointer to the corresponding full-text information in the information chain in each block of the index chain.

[0036] Furthermore, in a blockchain-based archive retrieval system, the retrieval index information module includes:

[0037] A search request receiving module, used to receive a search request from a user, wherein the search request includes a keyword or phrase that the user wishes to search for;

[0038] A preprocessing module, used for preprocessing the search request, wherein the preprocessing includes removing stop words and segmenting words;

[0039] In the module for screening relevant archives, the smart contract calculates the matching degree between the index information on the index chain and the user's search request through the archive relevance model;

[0040] The sorting module is used to calculate the results according to the matching degree and sort the index information items of all candidate archives through the index optimization model to ensure that the index information items of the most relevant archives are ranked first.

[0041] Furthermore, in a blockchain-based archive retrieval system, the module for returning full-text information includes:

[0042] A full-text information positioning module is used to locate the full-text information of the electronic archive according to the pointer pointing to the full-text information in the index information;

[0043] The hash value comparison verification module is used to compare the hash value of the archive content with the hash value stored on the blockchain through the smart contract;

[0044] The module returns after matching, and is used to return the full text information list of the archive to the user if the comparison is consistent.

[0045] By using blockchain and smart contracts, the storage and retrieval process of archives increases transparency and immutability. Every access and modification of archives will be recorded, reducing the risk of data tampering and forgery.

[0046] Storing index information and full-text information separately can speed up retrieval. Users can quickly obtain a list of archives most relevant to the search request without having to download and view the full-text information of each archive to determine relevance. Smart contracts can calculate the matching degree between search requests and archives based on complex algorithms to provide more accurate search results. In addition, smart contracts can ensure that the archive verification process is fair and transparent, increasing the credibility of search results. Users can find the required archives faster and have higher confidence in the authenticity and integrity of the archives. The solution also supports secure access to and use of archives, improving the overall user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 This is a schematic diagram of a blockchain-based archive retrieval method according to the present invention. DETAILED DESCRIPTION

[0048] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0049] One embodiment of the present invention provides a blockchain-based archive retrieval method, the method comprising:

[0050] The electronic files forming the archives are electronic archives, which are divided into index information and full-text information, and are placed on the index chain and the full-text information chain respectively;

[0051] Receive a user search request, calculate the matching degree between the search request and the archive entry through a smart contract, and return a list of index information of the archives most relevant to the user search request;

[0052] The full-text information of the electronic archive is located according to the index information, and the archive is verified through the smart contract. Once the verification is successful, a list of full-text archive information that is successfully verified and most relevant to the search criteria is returned to the user.

[0053] The working principle of the above technical solution is as follows: the electronic archive is split into two parts: index information and full-text information. The index information is stored on a special index chain, while the full-text information is stored on another full-text information chain. This split makes the retrieval process more efficient because the index information is usually smaller and more suitable for fast scanning and matching; receiving retrieval requests submitted by users, processing these requests through smart contracts, and calculating the degree of match between the retrieval request and the archive index entry according to a preset algorithm, and returning a sorted index information list. The archives in the list are the most relevant to the user's query, and the corresponding full-text information is located through the index information. Then, the smart contract is used to verify the authenticity and integrity of the archive, which may include verifying the archive's digital signature, timestamp or other information related to the security and trustworthiness of the archive; once the archive is verified, the system returns a list of full-text information of these archives to the user. This list contains the full-text information of the archives that best matches the user's search conditions and has been successfully verified.

[0054] The effects of the above technical solutions are as follows: by using blockchain and smart contracts, the storage and retrieval process of archives increases transparency and immutability, and each access and modification of archives will be recorded, reducing the risk of data tampering and forgery; storing index information and full-text information separately can speed up retrieval, and users can quickly obtain a list of archives most relevant to the retrieval request without having to download and view the full-text information of each archive to determine the relevance; smart contracts can calculate the matching degree between retrieval requests and archives based on complex algorithms, and provide more accurate retrieval results. In addition, smart contracts can ensure that the verification process of archives is fair and transparent, increasing the credibility of retrieval results. Users can find the required archives faster and have higher confidence in the authenticity and integrity of the archives. The solution also supports secure access to and use of archives, improving the overall user experience.

[0055] One embodiment of the present invention is a blockchain-based archive retrieval method, which forms an electronic file of the archive, namely an electronic archive, splits the electronic archive into index information and full-text information, and places them on the index chain and the full-text information chain respectively, including:

[0056] Forming electronic archives, extracting index information from each electronic archive, the index information including: the number of times the archive has been queried, the length of the electronic archive, the archive timestamp, the frequency of modification of the archive, the number of times the archive has been cited, the number of times the archive has been queried, the time of the most recent modification of the archive and the archive summary, and placing the index information on an index chain;

[0057] Calculate the hash value of the electronic file, place the full-text information of the electronic file together with the hash value on the information chain, and retain a pointer to the corresponding full-text information in the information chain in each block of the index chain.

[0058] The working principle of the above technical solution is as follows: extract key index information from each electronic archive, such as the number of times the archive has been queried, the size of the archive, the timestamp, the frequency of modification, the number of mutual references between archives, and the archive summary. This information helps to quickly evaluate the relevance and importance of the archive; place the extracted index information on the index chain, which is a blockchain structure specifically used to store index information, which can optimize retrieval performance because this information is usually small and used for fast screening and retrieval; calculate the hash value of each electronic archive, and store the full-text information and its hash value on another blockchain called the information chain. The hash value is used to verify the integrity and consistency of the archive; in each block of the index chain, retain a pointer to the corresponding full-text information in the information chain. In this way, once the relevant archive is found through the index information, its full-text content can be directly accessed through the pointer.

[0059] The effects of the above technical solution are: improving retrieval efficiency: by storing index information and full-text information separately, the system can process retrieval requests faster, because the index chain is smaller and easier to search, users can quickly find relevant archive indexes, and then decide whether to access full-text information; enhancing data security, using hash values ​​and blockchain technology can ensure the immutability and consistency of archive data, hash values ​​as proof of archive integrity, any modification to the archive will cause the hash value to change, so that unauthorized changes or tampering can be easily detected. This increases the trust of data, especially in areas with high legal and compliance requirements; ensuring data traceability and transparency, storing archive index information and full-text hash values ​​on the blockchain makes every archive access and modification recorded and traceable, which facilitates audits and compliance checks and ensures transparency of operations; through blockchain technology, electronic archives can be managed more effectively. The separation design of index chain and information chain not only optimizes the retrieval and storage efficiency of data, but also can flexibly expand system capacity as needed, so that users can get faster responses and higher data security guarantees when retrieving archives. At the same time, the most relevant archive indexes can be quickly screened out through the index chain, and the full-text information can be directly accessed through pointers, making the whole process more efficient and user-friendly; this technical solution provides a secure and efficient archive management system by combining the immutability of blockchain, hash-verified data integrity, and structured storage of indexes and information.

[0060] One embodiment of the present invention is a blockchain-based archive retrieval method, which receives a user retrieval request, calculates the matching degree between the retrieval request and the archive item through a smart contract, and returns a list of archives most relevant to the user retrieval request, including:

[0061] Receiving a search request from a user, wherein the search request includes a keyword or phrase that the user wishes to search for;

[0062] Preprocessing the search request, wherein the preprocessing includes removing stop words and segmenting words;

[0063] The smart contract calculates the matching degree between the index information on the index chain and the user's search request through the archive relevance model;

[0064] According to the matching degree calculation results, the index information entries of all candidate archives are sorted through the index optimization model to ensure that the index information entries of the most relevant archives are ranked first.

[0065] The working principle of the above technical solution is as follows: the user submits a search request containing keywords or phrases, and these keywords are the content that the user hopes to find in the archives; the user's search request is pre-processed, including removing stop words (i.e. common words that do not provide useful information in the search, such as "and", "is", etc.) and performing word segmentation (splitting long sentences or phrases into smaller units that are easier to process); the smart contract uses the archive relevance model to calculate the degree of match between each archive index information on the index chain and the user's search request. The relevance model may be based on factors such as the frequency and location of keyword occurrence, and the update time of the archive; based on the calculation result of the matching degree, the smart contract sorts the index information items of all candidate archives through the index optimization model to ensure that the most relevant archive index information items are ranked first.

[0066] The effects of the above technical solution are: improving retrieval accuracy. By preprocessing retrieval requests and using advanced relevance models, the system can more accurately identify user needs and return the archives most relevant to the user's query; the automated processing of smart contracts reduces the need for manual intervention and speeds up the retrieval process. At the same time, the index information is optimized and sorted so that users can find the most relevant information more quickly; users can expect faster response times and more accurate search results, thereby improving overall satisfaction and efficiency; since the entire process is carried out on the blockchain, the immutability and transparency of the data are ensured, and the trust of the system is increased. The solution provides users with a fast, accurate and secure electronic archive retrieval experience by combining blockchain technology, smart contracts and advanced search algorithms.

[0067] An embodiment of the present invention provides a blockchain-based archive retrieval method, wherein the archive relevance model includes:

[0068] The archive correlation model is:

[0069]

[0070] Among them, W is the relevance of the archive, f is the frequency of keywords in the summary of the electronic archive index information, t0 is the time when the retrieval request is received, F is the user feedback score, t is the timestamp of the archive, L is the length of the electronic archive, and α, β, λ and θ are adjustment parameters.

[0071] The working principle and effect of the above technical solution are: the frequency of keyword occurrence f reflects the importance of the file; timeliness It ensures that the most recent archives are more likely to be retrieved. These two factors reflect the archives that best match the user's search and the user's preference for the latest archives, helping users quickly find the archives that best match the user's search and the current popular archives, increasing the timeliness and attractiveness of the search results; user feedback (such as likes and number of comments) is an intuitive indicator for measuring archive quality and user satisfaction. It provides users with a direct evaluation of the value of the archive. By giving priority to displaying archives with high feedback scores, it can better meet user expectations and needs, increase user participation and the interactivity of the retrieval system; document length affects the level of detail of information and user reading experience. Longer documents may contain more comprehensive information, but may also cause users to feel burdened. Appropriate length assessment can help balance the richness and readability of information; ensure that the search results can provide sufficient information without overwhelming users due to excessively long documents. This design aims to improve user reading experience and satisfaction; adjustment parameters allow system administrators to adjust the weights of user feedback scores and document length factors according to specific application scenarios or the needs of user groups. This flexibility It is to cope with different types of information needs and retrieval environments, and provide a mechanism. By adjusting these parameters, the system can better adapt to the specific preferences of different user groups and improve the applicability and efficiency of the retrieval system; comprehensive consideration of user feedback and document length provides users with a more personalized retrieval experience, thereby greatly improving user satisfaction; by dynamically adjusting parameters, the system can flexibly respond to the needs of different situations, so that the retrieval results are more in line with the actual needs of users; through refined weight calculations, the system can more accurately identify and provide high-quality, highly relevant information, reduce information overload, and improve retrieval efficiency; by comprehensively considering the key factors in the information retrieval process, a more accurate, efficient and user-satisfying retrieval experience is achieved.

[0072] One embodiment of the present invention is a blockchain-based archive retrieval method, wherein the index optimization model includes:

[0073] The index optimization model is:

[0074]

[0075] Among them, P represents the sorting priority, λ, μ and v are the weight coefficients that control the contribution of each part and can be adjusted according to the retrieval requirements.q and w s They represent the number of file queries and the weight of user preference, q and s represent the number of queries and user preference statistics, Q max and S max are the maximum expected values ​​of q and s, respectively. T represents the freshness of the document, that is, the time from the last modification to now. T max is the maximum value of T and is used to normalize T. MR is a combined factor that takes into account the modification frequency and the number of citations. max is its maximum expected value, and κ is a nonlinear coefficient used to adjust the influence of MR. Optionally, MR can be obtained by the formula MR=w M ·M+w R R calculation, where M is the modification frequency, R is the number of times the file is referenced, and W M and W R It is a pre-set weight that is adjusted according to the importance of modification frequency and citation count on priority.

[0076] After finding a series of related archives through the archive correlation model, there may be many electronic archives with similar weights. In this case, the index optimization formula gives different priorities to these archives by considering the number of archive queries and size, and determines their display order. The introduction of the index optimization formula is to further optimize the presentation of retrieval results and system performance after retrieving related archives through the archive correlation model. It gives archives more refined priorities by considering the query frequency of archives and user preferences, thereby improving the user's retrieval experience.

[0077] The working principle and effect of the above technical solution are as follows: the two factors of archive query times and user preferences measure the relevance of information to user queries and the personalized preferences of users, respectively. By weighting them, it can be ensured that the system takes into account both the specific needs of users and the historical preferences of users, increasing the personalization and relevance of retrieval results, so that users can obtain information that better suits their needs and interests; by normalizing q and s, it can be ensured that the comparison between different archives is fair, regardless of the range of their original values, so that all archives are at the same level when scoring, which is convenient for comparison and sorting, and improves the consistency and accuracy of the scoring system; considering that users usually prefer updated information, the priority is adjusted by the freshness of the document to ensure that users can more easily obtain the latest information, improve user experience and satisfaction; the combined factor of modification frequency and citation number, a document that is frequently modified or widely cited usually means that it is highly active or important, by emphasizing these widely recognized or frequently updated information, which improves the quality and practicality of the retrieval results; by introducing a nonlinear coefficient to adjust the impact of (MR), the contribution of modification frequency and citation count to the final score can be more finely controlled; allowing the system to more flexibly reflect the characteristics of different types of documents, such as paying attention to frequently updated documents or widely cited documents, increasing the flexibility and accuracy of scoring. By considering multi-dimensional factors and combining their contributions, this formula can more accurately calculate the priority of archives, thereby providing retrieval results that are more in line with user needs and preferences; ensuring that users can quickly find the most relevant, fresh and high-quality information, thereby greatly improving user satisfaction and trust in the retrieval system; by adjusting different weight coefficients and nonlinear coefficients, the focus of priority scoring can be adjusted according to different application scenarios or the needs of user groups.

[0078] An embodiment of the present invention is a blockchain-based archive retrieval method, which locates the full-text information of electronic archives according to index information, verifies the archives through smart contracts, and returns a list of archives full-text information that are successfully verified and most relevant to the search conditions to the user after the verification is passed, including:

[0079] Locate the full-text information of the electronic archive according to the pointer to the full-text information in the index information;

[0080] Compare the hash value of the archive content with the hash value stored on the blockchain through the smart contract;

[0081] If the comparison is consistent, the full text information list of the archive is returned to the user.

[0082] The working principle and effect of the above technical solution are as follows: the index information of the electronic archive contains a pointer to the full-text information (for example, a URL or file path). When a specific electronic archive needs to be found, the system first queries this index information to determine the location where the full-text information is stored; hash value comparison is performed using smart contracts. When each electronic archive is added to the system, the hash value of its content is recorded on the blockchain through a smart contract. When a request is made to access a specific archive, the system will compare the current hash value of the archive with the hash value stored on the blockchain; the smart contract automatically executes the comparison process to ensure that the content of the archive being accessed has not been tampered with or changed. If the hash value of the current archive is consistent with the hash value recorded on the blockchain, it means that the archive is original and has not been modified; once the verification process is successful, that is, the hash values ​​match, the system will return a list of the archive's full-text information to the user, and the user can access and view the archive content.

[0083] The effects of the above technical solution are: by hashing the archive content and storing the hash value on the blockchain, the data is ensured to be tamper-proof. Even if the database is attacked, the hash value comparison mechanism can detect whether the data has been illegally modified; the data on the blockchain is highly consistent and persistent. Once the data is written to the blockchain, it is almost impossible to modify or delete it, thus ensuring the reliability of the electronic archives; each storage or access operation of the archive is recorded on the blockchain through a smart contract, providing a complete audit trail function. This makes any access and operation to the electronic archives transparent and easy to track; the process of automatically comparing and verifying hash values ​​using smart contracts reduces manual intervention and improves processing speed and efficiency.

[0084] An embodiment of the present invention provides a blockchain-based archive retrieval system, the system comprising:

[0085] Generate an electronic archive module, which is used to form an electronic file of the archive, namely, an electronic archive, split the electronic archive into index information and full-text information, and place them on the index chain and the full-text information chain respectively;

[0086] A retrieval index information module is used to receive a user retrieval request, calculate the matching degree between the retrieval request and the archive entry through a smart contract, and return a list of index information of archives most relevant to the user's retrieval request;

[0087] The full-text information return module is used to locate the full-text information of the electronic archive according to the index information, and verify the archive through the smart contract. If the verification is successful, a list of full-text information of the archives that are successfully verified and most relevant to the search criteria is returned to the user.

[0088] The working principle of the above technical solution is as follows: the electronic archive is split into two parts: index information and full-text information. The index information is stored on a special index chain, while the full-text information is stored on another full-text information chain. This split makes the retrieval process more efficient because the index information is usually smaller and more suitable for fast scanning and matching; receiving retrieval requests submitted by users, processing these requests through smart contracts, and calculating the degree of match between the retrieval request and the archive index entry according to a preset algorithm, and returning a sorted index information list. The archives in the list are the most relevant to the user's query, and the corresponding full-text information is located through the index information. Then, the smart contract is used to verify the authenticity and integrity of the archive, which may include verifying the archive's digital signature, timestamp or other information related to the security and trustworthiness of the archive; once the archive is verified, the system returns a list of full-text information of these archives to the user. This list contains the full-text information of the archives that best matches the user's search conditions and has been successfully verified.

[0089] The effects of the above technical solution are: by using blockchain and smart contracts, the storage and retrieval process of archives increases transparency and immutability, and each access and modification of archives will be recorded, reducing the risk of data tampering and forgery;

[0090] Storing index information and full-text information separately can speed up retrieval. Users can quickly obtain a list of archives most relevant to the search request without having to download and view the full-text information of each archive to determine relevance. Smart contracts can calculate the matching degree between search requests and archives based on complex algorithms to provide more accurate search results. In addition, smart contracts can ensure that the archive verification process is fair and transparent, increasing the credibility of search results. Users can find the required archives faster and have higher confidence in the authenticity and integrity of the archives. The solution also supports secure access to and use of archives, improving the overall user experience.

[0091] An embodiment of the present invention is a blockchain-based archive retrieval system, wherein the electronic archive generation module includes:

[0092] Generate an index chain module, which is used to form an electronic archive, extract index information from each electronic archive, the index information includes: the number of times the archive is queried, the length of the electronic archive, the archive timestamp, the modification frequency of the archive, the number of times the archive is cited, the number of times the archive is queried, the time of the most recent modification of the archive and the archive summary, and place the index information on the index chain;

[0093] Generate an information chain module, which is used to calculate the hash value of the electronic file, place the full-text information of the electronic file and the hash value together on the information chain, and retain a pointer to the corresponding full-text information in the information chain in each block of the index chain.

[0094] The working principle of the above technical solution is as follows: extract key index information from each electronic archive, such as the number of times the archive has been queried, the size of the archive, the timestamp, the frequency of modification, the number of mutual references between archives, and the archive summary. This information helps to quickly evaluate the relevance and importance of the archive; place the extracted index information on the index chain, which is a blockchain structure specifically used to store index information, which can optimize retrieval performance because this information is usually small and used for fast screening and retrieval; calculate the hash value of each electronic archive, and store the full-text information and its hash value on another blockchain called the information chain. The hash value is used to verify the integrity and consistency of the archive; in each block of the index chain, retain a pointer to the corresponding full-text information in the information chain. In this way, once the relevant archive is found through the index information, its full-text content can be directly accessed through the pointer.

[0095] The effects of the above technical solution are: improving retrieval efficiency: by storing index information and full-text information separately, the system can process retrieval requests faster, because the index chain is smaller and easier to search, users can quickly find relevant archive indexes, and then decide whether to access full-text information; enhancing data security, using hash values ​​and blockchain technology can ensure the immutability and consistency of archive data, hash values ​​as proof of archive integrity, any modification to the archive will cause the hash value to change, so that unauthorized changes or tampering can be easily detected. This increases the trust of data, especially in areas with high legal and compliance requirements; ensuring data traceability and transparency, storing archive index information and full-text hash values ​​on the blockchain makes every archive access and modification recorded and traceable, which facilitates audits and compliance checks and ensures transparency of operations; through blockchain technology, electronic archives can be managed more effectively. The separation design of index chain and information chain not only optimizes the retrieval and storage efficiency of data, but also can flexibly expand system capacity as needed, so that users can get faster responses and higher data security guarantees when retrieving archives. At the same time, the most relevant archive indexes can be quickly screened out through the index chain, and the full-text information can be directly accessed through pointers, making the whole process more efficient and user-friendly; this technical solution provides a secure and efficient archive management system by combining the immutability of blockchain, hash-verified data integrity, and structured storage of indexes and information.

[0096] An embodiment of the present invention is a blockchain-based archive retrieval system, wherein the retrieval index information module includes:

[0097] A search request receiving module, used to receive a search request from a user, wherein the search request includes a keyword or phrase that the user wishes to search for;

[0098] A preprocessing module, used for preprocessing the search request, wherein the preprocessing includes removing stop words and segmenting words;

[0099] In the module for screening relevant archives, the smart contract calculates the matching degree between the index information on the index chain and the user's search request through the archive relevance model;

[0100] The sorting module is used to calculate the results according to the matching degree and sort the index information items of all candidate archives through the index optimization model to ensure that the index information items of the most relevant archives are ranked first.

[0101] The working principle of the above technical solution is as follows: the user submits a search request containing keywords or phrases, and these keywords are the content that the user hopes to find in the archives; the user's search request is pre-processed, including removing stop words (i.e. common words that do not provide useful information in the search, such as "and", "is", etc.) and performing word segmentation (splitting long sentences or phrases into smaller units that are easier to process); the smart contract uses the archive relevance model to calculate the degree of match between each archive index information on the index chain and the user's search request. The relevance model may be based on factors such as the frequency and location of keyword occurrence, and the update time of the archive; based on the calculation result of the matching degree, the smart contract sorts the index information items of all candidate archives through the index optimization model to ensure that the most relevant archive index information items are ranked first.

[0102] The effects of the above technical solution are: improving retrieval accuracy. By preprocessing retrieval requests and using advanced relevance models, the system can more accurately identify user needs and return the archives most relevant to the user's query; the automated processing of smart contracts reduces the need for manual intervention and speeds up the retrieval process. At the same time, the index information is optimized and sorted so that users can find the most relevant information more quickly; users can expect faster response times and more accurate search results, thereby improving overall satisfaction and efficiency; since the entire process is carried out on the blockchain, the immutability and transparency of the data are ensured, and the trust of the system is increased. The solution provides users with a fast, accurate and secure electronic archive retrieval experience by combining blockchain technology, smart contracts and advanced search algorithms.

[0103] An embodiment of the present invention is a blockchain-based archive retrieval system, wherein the module for returning full-text information includes:

[0104] A full-text information positioning module is used to locate the full-text information of the electronic archive according to the pointer pointing to the full-text information in the index information;

[0105] The hash value comparison verification module is used to compare the hash value of the archive content with the hash value stored on the blockchain through the smart contract;

[0106] The module returns after matching, and is used to return the full text information list of the archive to the user if the comparison is consistent.

[0107] The working principle and effect of the above technical solution are as follows: the index information of the electronic archive contains a pointer to the full-text information (for example, a URL or file path). When a specific electronic archive needs to be found, the system first queries this index information to determine the location where the full-text information is stored; hash value comparison is performed using smart contracts. When each electronic archive is added to the system, the hash value of its content is recorded on the blockchain through a smart contract. When a request is made to access a specific archive, the system will compare the current hash value of the archive with the hash value stored on the blockchain; the smart contract automatically executes the comparison process to ensure that the content of the archive being accessed has not been tampered with or changed. If the hash value of the current archive is consistent with the hash value recorded on the blockchain, it means that the archive is original and has not been modified; once the verification process is successful, that is, the hash values ​​match, the system will return a list of the archive's full-text information to the user, and the user can access and view the archive content.

[0108] The effects of the above technical solution are: by hashing the archive content and storing the hash value on the blockchain, the data is ensured to be tamper-proof. Even if the database is attacked, the hash value comparison mechanism can detect whether the data has been illegally modified; the data on the blockchain is highly consistent and persistent. Once the data is written to the blockchain, it is almost impossible to modify or delete it, thus ensuring the reliability of the electronic archives; each storage or access operation of the archive is recorded on the blockchain through a smart contract, providing a complete audit trail function. This makes any access and operation to the electronic archives transparent and easy to track; the process of automatically comparing and verifying hash values ​​using smart contracts reduces manual intervention and improves processing speed and efficiency.

[0109] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A blockchain-based archive retrieval method, characterized in that: The method comprises: The electronic files forming the archives are electronic archives, which are divided into index information and full-text information, and are placed on the index chain and the full-text information chain respectively; Receive a user search request, calculate the matching degree between the search request and the archive entry through a smart contract, and return a list of index information of the archives most relevant to the user search request; The full-text information of the electronic archive is located according to the index information, and the archive is verified through the smart contract. Once the verification is successful, a list of full-text archive information that is successfully verified and most relevant to the search criteria is returned to the user.

2. According to the blockchain-based archive retrieval method of claim 1, it is characterized in that: The electronic files that form the archives are electronic archives, which are divided into index information and full-text information, and are placed on the index chain and the full-text information chain respectively, including: Forming electronic archives, extracting index information from each electronic archive, the index information including: the number of times the archive has been queried, the length of the electronic archive, the archive timestamp, the frequency of modification of the archive, the number of times the archive has been cited, the number of times the archive has been queried, the time of the most recent modification of the archive and the archive summary, and placing the index information on an index chain; Calculate the hash value of the electronic file, place the full-text information of the electronic file together with the hash value on the information chain, and retain a pointer to the corresponding full-text information in the information chain in each block of the index chain.

3. According to the blockchain-based archive retrieval method of claim 1, it is characterized in that: Receive a user search request, calculate the matching degree between the search request and the archive item through the smart contract, and return a list of archives most relevant to the user's search request, including: Receiving a search request from a user, wherein the search request includes a keyword or phrase that the user wishes to search for; Preprocessing the search request, wherein the preprocessing includes removing stop words and segmenting words; The smart contract calculates the matching degree between the index information on the index chain and the user's search request through the archive relevance model; According to the matching degree calculation results, the index information entries of all candidate archives are sorted through the index optimization model to ensure that the index information entries of the most relevant archives are ranked first.

4. According to the blockchain-based archive retrieval method of claim 3, it is characterized in that: The archive correlation model includes: The archive correlation model is: Among them, W is the relevance of the archive, f is the frequency of keywords in the abstract of the electronic archive index information, t0 is the time when the retrieval request is received, t is the timestamp of the archive, L is the length of the electronic archive, and α, β and θ are weight coefficients.

5. According to the blockchain-based archive retrieval method of claim 3, it is characterized in that: The index optimization model includes: The index optimization model is: Among them, P represents the sorting priority, λ, μ and v are the weight coefficients that control the contribution of each part and can be adjusted according to the retrieval requirements. q and w s They represent the number of file queries and the weight of user preference, q and s represent the number of queries and user preference statistics, Q max and S max are the maximum expected values ​​of q and s respectively, T represents the freshness of the document, that is, the time from the last modification to now, T max It is the maximum value of T and is used to normalize T. MR is a combined factor that takes into account the modification frequency and the number of citations. max is its maximum expected value, and κ is a nonlinear coefficient used to adjust the influence of MR.

6. According to the blockchain-based archive retrieval method of claim 1, it is characterized in that: The full-text information of the electronic archive is located according to the index information, and the archive is verified through the smart contract. If the verification is successful, a list of full-text information of the archive that is successfully verified and most relevant to the search criteria is returned to the user, including: Locate the full-text information of the electronic archive according to the pointer to the full-text information in the index information; Compare the hash value of the archive content with the hash value stored on the blockchain through the smart contract; If the comparison is consistent, the full text information list of the archive is returned to the user.

7. A blockchain-based archive retrieval system, characterized in that: The system comprises: Generate an electronic archive module, which is used to form an electronic file of the archive, namely, an electronic archive, split the electronic archive into index information and full-text information, and place them on the index chain and the full-text information chain respectively; A retrieval index information module is used to receive a user retrieval request, calculate the matching degree between the retrieval request and the archive entry through a smart contract, and return a list of index information of archives most relevant to the user's retrieval request; The full-text information return module is used to locate the full-text information of the electronic archive according to the index information, and verify the archive through the smart contract. If the verification is successful, a list of full-text information of the archives that are successfully verified and most relevant to the search criteria is returned to the user.

8. According to claim 7, a blockchain-based archive retrieval system is characterized in that: The electronic archive generation module includes: Generate an index chain module, which is used to form an electronic archive, extract index information from each electronic archive, the index information includes: the number of times the archive is queried, the length of the electronic archive, the archive timestamp, the modification frequency of the archive, the number of times the archive is cited, the number of times the archive is queried, the time of the most recent modification of the archive and the archive summary, and place the index information on the index chain; Generate an information chain module, which is used to calculate the hash value of the electronic file, place the full-text information of the electronic file and the hash value together on the information chain, and retain a pointer to the corresponding full-text information in the information chain in each block of the index chain.

9. According to claim 7, a blockchain-based archive retrieval system is characterized in that: The retrieval index information module includes: A search request receiving module, used to receive a search request from a user, wherein the search request includes a keyword or phrase that the user wishes to search for; A preprocessing module, used for preprocessing the search request, wherein the preprocessing includes removing stop words and segmenting words; In the module for screening relevant archives, the smart contract calculates the matching degree between the index information on the index chain and the user's search request through the archive relevance model; The sorting module is used to calculate the results according to the matching degree and sort the index information items of all candidate archives through the index optimization model to ensure that the index information items of the most relevant archives are ranked first.

10. According to claim 7, a blockchain-based archive retrieval system is characterized in that: The module for returning full-text information includes: A full-text information positioning module is used to locate the full-text information of the electronic archive according to the pointer pointing to the full-text information in the index information; The hash value comparison verification module is used to compare the hash value of the archive content with the hash value stored on the blockchain through the smart contract; The module returns after matching, and is used to return the full text information list of the archive to the user if the comparison is consistent.