Data integrity protection method based on hash function and block chain technology

By using hash function and blockchain technology in property rights data, we calculate the characteristic vocabulary correlation between the text to be tested and the property rights text and judge the infringement indicators, the problem of low accuracy of infringement judgment in the existing technology is solved, and the accuracy and security of data integrity protection are improved.

CN120105489AActive Publication Date: 2025-06-06JINJING (HAINAN) TECH DEV CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510178144.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-06
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

The prior art does not represent the validity of property rights in property rights data, resulting in a low accuracy of infringement judgment on property rights data.

Method used

The data integrity protection method based on hash function and blockchain technology is adopted. By obtaining the characteristic vocabulary of the text to be tested and the property rights text stored on the blockchain, the overall subject correlation degree of the characteristic vocabulary is calculated, and based on this, the infringement indicators of the property rights text to be tested are obtained to protect the data integrity of the blockchain.

Benefits of technology

It improves the accuracy of judgment on document infringement, avoids malicious modifications to property rights information by the text to be tested before being put on the chain, and ensures the security and reliability of user property rights.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105489A_ABST
    Figure CN120105489A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data integrity protection, in particular to a data integrity protection method based on a hash function and a block chain technology. The method comprises the following steps: determining feature vocabularies in a property right text and a to-be-tested text; according to the character number and the vocabulary number of feature vocabularies in a to-be-tested text appearing in a reproduction paragraph of each feature vocabulary in the property right text, and the similar condition of the entity relationship between the clause where the feature vocabularies of the to-be-tested text appearing in the reproduction paragraph are located and the other clauses, the character vocabularies in the to-be-tested text are obtained; obtaining the overall subject association degree of each feature vocabulary of the to-be-tested text relative to the property right text; and acquiring infringement indexes of the to-be-tested text to the property right text according to the overall subject association degree difference and the text difference of the feature vocabularies of the to-be-tested text and the property right text, thereby protecting the data integrity of the block chain. According to the invention, the accuracy of infringement judgment on the property right data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of protecting data integrity, and in particular to a data integrity protection method based on hash functions and blockchain technology. Background Art

[0002] When protecting data in the property registration system, blockchain technology is usually used to merge data blocks to form an encrypted chain data structure. The characteristics of decentralized trust are used to improve the traditional method of single point failure or excessive load on the central node leading to data loss, damage and leakage in the system. At the same time, blockchain has the characteristic of being tamper-proof, so the application of blockchain technology in the property registration system can significantly improve the security and reliability of property data.

[0003] Before the property data is stored in the blockchain, the property information may be maliciously copied or modified, resulting in the property data possibly infringing the user's property rights stored in the blockchain. Existing methods usually use the word frequency-inverse document frequency algorithm to determine whether the property data is infringing. However, in the process of selecting key words by the word frequency-inverse document frequency algorithm, the frequency of some words representing the validity of property rights is not prominent, thereby reducing the accuracy of the infringement judgment of the document. Summary of the invention

[0004] In order to solve the technical problem that the words representing the validity of property rights in property rights data are not prominent, resulting in a low accuracy rate in judging infringement of property rights data, the purpose of the present invention is to provide a data integrity protection method based on hash function and blockchain technology. The technical solution adopted is as follows:

[0005] The present invention proposes a data integrity protection method based on hash function and blockchain technology, the method comprising:

[0006] Obtain the text to be tested and the property rights text stored on the blockchain, and record the text to be tested and the property rights text as analysis text; divide the analysis text into different sentences, and segment the sentences to obtain segmented words;

[0007] According to the number of occurrences of each participle of each sentence in the analyzed text and the other participles in the other sentences, characteristic words in the analyzed text are obtained;

[0008] Determine the recurring paragraphs of each characteristic word in the test text in the property text; obtain the overall subject relevance of each characteristic word in the test text to the property text based on the number of characters and the number of words in each recurring paragraph of the characteristic word in the test text, and the similarity between the entity relationship between the sentence in which the characteristic word in the test text appears in the recurring paragraph and the other sentences in the property text;

[0009] According to the overall subject correlation difference and text difference between the characteristic vocabulary of the text to be tested and the property text, the infringement index of the text to be tested on the property text is obtained; based on the infringement index, the data integrity of the blockchain is protected.

[0010] Furthermore, the acquisition of characteristic words in the analysis text includes:

[0011] Arrange the participles of each sentence in the analyzed text in order of position to obtain the participle sequence of the corresponding sentence;

[0012] For each sentence segmentation sequence of the analyzed text, a segmentation in the segmentation sequence is randomly selected and recorded as a target segmentation, and the target segmentation and each segmentation after it constitute a segmentation pair of the target segmentation; in the analyzed text, the proportion of the number of sentences of each segmentation pair containing the target segmentation in the number of sentences containing the target segmentation is used as the characteristic degree of each segmentation pair of the target segmentation; the segmentation pair corresponding to the maximum value of the characteristic degree of the target segmentation is selected and recorded as the initial characteristic vocabulary of the target segmentation;

[0013] The initial characteristic words of all the word segments in the analyzed text are recorded as the characteristic words in the analyzed text.

[0014] Furthermore, the step of obtaining the overall subject relevance of each characteristic word of the text to be tested relative to the property text includes:

[0015] According to the similarity of the entity relationship between the clause where each characteristic word is located and the other clauses in the analyzed text, and the distribution position of the other clauses in the analyzed text, the lexical importance index of each characteristic word in the analyzed text is obtained;

[0016] The overall subject relevance of each characteristic word in the test text to the property text is obtained based on the number of characters and the number of words in each characteristic word's recurring paragraph in the property text, as well as the lexical importance index of the characteristic words in the test text that appear in the recurring paragraph.

[0017] Furthermore, the acquisition of the vocabulary importance index of each characteristic vocabulary in the analysis text includes:

[0018] Entity relations are extracted for each sentence of the analyzed text to obtain the SPO triples of the corresponding sentences; the proportion of the number of characters of the SPO triples in the total number of characters of the corresponding sentences is recorded as the sentence ratio of the SPO triples;

[0019] A characteristic word in the analysis text is recorded as an example word, and a similar triplet of the example word is selected from the SPO triples of each sentence in the analysis text except the sentence where the example word is located, and the similar triplet is equal to the first element and the third element of any SPO triplet of the sentence where the example word is located; the average of the sentence ratios of all similar triples of the example word is used as the entity similarity index of the example word;

[0020] Number the sentences in the analyzed text, and calculate the mean of the sequence numbers of the sentences where all similar triples of the sample words are located, as the position distribution indicator of the sample words;

[0021] According to the number of similar triples of the example vocabulary, the entity similarity index and the position distribution index, the vocabulary importance index of the example word segmentation is obtained; the number of similar triples of the example vocabulary and the entity similarity index are both positively correlated with the vocabulary importance index, and the position distribution index is negatively correlated with the vocabulary importance index.

[0022] Furthermore, the step of obtaining the overall subject relevance of each characteristic word in the test text to the property text includes:

[0023] A characteristic word in the test text is randomly selected as the target word, and the ratio of the total number of characters of the characteristic word in the test text that appears in each recurring paragraph of the target word in the property text to the total number of characters in the recurring paragraph is taken as the character repetition degree of the target word in each recurring paragraph of the property text;

[0024] Select characteristic words that appear in each recurring paragraph of the target word in the property text in the test text, and record them as the focus words of the recurring paragraph; take the ratio of the mean value of the lexical importance index of the focus words in each recurring paragraph of the target word in the property text to the maximum value of the lexical importance index of the characteristic words in the property text as the text importance of the target word in each recurring paragraph of the property text;

[0025] According to the character repetition degree, the text importance, and the number of words in which the characteristic words in the test text appear in each recurring paragraph of the target word in the property text, the local subject relevance of the target word in each recurring paragraph of the property text is obtained;

[0026] The average of the local subject relevance of the target vocabulary in all recurring paragraphs of the property text is used as the overall subject relevance of the target vocabulary relative to the property text.

[0027] Furthermore, the step of obtaining the infringement index of the tested text on the property text includes:

[0028] The ratio of the frequency of occurrence of each characteristic word of the test text in the test text to the average frequency of occurrence of the characteristic word in all property texts on the blockchain is used as the block association index of each characteristic word of the test text; the block association index of each characteristic word of the test text is negatively correlated, and the product of the mapping result and the overall subject association degree is used as the recognition support index of each characteristic word of the test text;

[0029] Perform hash processing on the text to be tested to obtain the hash value of the text to be tested; obtain the difference sequence number, the hash values ​​of the text to be tested and the property text are different in the characters of each difference sequence number; perform word vector conversion on the feature words corresponding to the characters of the hash value of the analysis text at each difference sequence number to obtain the feature vector of the hash value of the analysis text at each difference sequence number; take the average value of the distance between the feature vectors of the hash values ​​of the text to be tested and the property text at the same difference sequence number as the semantic difference value between the text to be tested and the analysis text;

[0030] The ratio of the number of difference serial numbers between the text to be tested and the property text to the number of target serial numbers is used as the hash difference value between the text to be tested and the property text; based on the hash difference value and the semantic difference value, the infringement index of the text to be tested on the property text is obtained.

[0031] Furthermore, the hash processing is performed on the text to be tested to obtain the hash value of the text to be tested, including:

[0032] The recognition support indicators of the characteristic words of the text to be tested are arranged in the order of the characteristic words' positions to obtain a recognition sequence of the text to be tested; a one-dimensional DCT transformation is performed on the recognition sequence to obtain a DCT coefficient of each element in the recognition sequence;

[0033] Calculate the mean of the DCT coefficients of all elements in the recognition sequence as the classification coefficient; set the encoding of the feature vocabulary corresponding to the elements in the recognition sequence whose DCT coefficient is greater than the classification coefficient to 1, and set the encoding of the feature vocabulary corresponding to the elements whose DCT coefficient is less than or equal to the classification coefficient to 0; arrange the encoding of the feature vocabulary of the text to be tested in the order of the position of the feature vocabulary to form a string as the hash value of the text to be tested.

[0034] Furthermore, protecting the data integrity of the blockchain based on the infringement indicator includes:

[0035] Determine whether the infringement indicators of the text to be tested for all property texts are less than the preset infringement threshold. If so, store the text to be tested on the node of the blockchain. Otherwise, the text to be tested cannot be stored on the node of the blockchain.

[0036] Furthermore, each characteristic word in the text to be tested contained in the recurring paragraph corresponds to two participles in the participle pair.

[0037] Furthermore, the number of target serial numbers is equal to the minimum value of the total number of characters of the hash values ​​of the text to be tested and the property rights text.

[0038] The present invention has the following beneficial effects:

[0039] In an embodiment of the present invention, there are professional terms and property rights words such as device structure design in the analysis text. When the sentence segmentation in the analysis text is performed, the property rights words cannot be accurately segmented. The possibility of different word segments constituting property rights words is measured by the number of occurrences of different word segments in each sentence in other sentences, and the characteristic vocabulary of the analysis text is obtained; the number of characters and the number of words in which the characteristic words in the test text appear in the recurring paragraphs of each characteristic word in the property rights text present the possibility of the main content of the test text appearing in the recurring paragraphs of the characteristic words, and the relationship between the entity of the sentence in the property rights text where the characteristic words of the test text appearing in the recurring paragraph are compared. Similar situations are presented, and the possibility that the main content in the text to be tested is the main content in the property text is presented. The two factors are combined to judge the degree of correlation between the characteristic vocabulary and the main content of the property text, and the overall subject correlation is obtained; through the overall subject correlation difference and text difference of the characteristic vocabulary of the text to be tested and the property text, the information similarity between the text to be tested and the property text is analyzed, and the possibility of the text to be tested infringing the property text is analyzed, and the infringement index is obtained to improve the accuracy of the infringement judgment of the document; based on the infringement index, it is decided whether to store the text to be tested on the blockchain, so as to avoid malicious modification of the property information before the text to be tested is uploaded to the chain, which may infringe the user's property rights. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0041] Figure 1 A flowchart of a method for protecting data integrity based on hash functions and blockchain technology provided by an embodiment of the present invention;

[0042] Figure 2 A flowchart of a method for obtaining the overall subject relevance provided by an embodiment of the present invention;

[0043] Figure 3 A flowchart of a method for obtaining infringement indicators provided by an embodiment of the present invention;

[0044] Figure 4A schematic diagram of a computer device of a data integrity protection device based on hash function and blockchain technology provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0045] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the data integrity protection method based on hash function and blockchain technology proposed by the present invention, its specific implementation method, structure, features and effects are described in detail as follows in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.

[0046] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0047] The specific scheme of the data integrity protection method based on hash function and blockchain technology provided by the present invention is described in detail below with reference to the accompanying drawings.

[0048] Embodiment 1:

[0049] This invention proposes a data integrity protection method based on hash function and blockchain technology, please refer to Figure 1 , which shows a flow chart of the steps of a data integrity protection method based on hash function and blockchain technology provided by an embodiment of the present invention, the method comprising:

[0050] Step S1: Obtain the text to be tested and the property rights text stored on the blockchain, and record the text to be tested and the property rights text as analysis texts; divide the analysis text into different sentences, and segment the sentences to obtain segmented words.

[0051] Specifically, the property registration text that needs to be stored on the blockchain is recorded as the test text, and the property registration text that has been stored on the node of the blockchain is recorded as the property text; wherein the blockchain stores multiple property texts. In order to analyze the infringement of the property text by the test text, the sentences in the analysis text are segmented.

[0052] In the embodiment of the present invention, the segmentation method is: first converting the analysis text into a text document format; then dividing the sentences in the analysis text based on punctuation marks to obtain different sentences, that is, recording the sentences between two adjacent punctuation marks as one sentence; finally using the Jieba word segmentation algorithm to segment each sentence to obtain multiple word segments. The Jieba word segmentation algorithm is a well-known technology for those skilled in the art and will not be described in detail here.

[0053] Step S2: Acquire characteristic words in the analysis text according to the number of occurrences of each participle of each sentence in the analysis text and the other participles in the other sentences.

[0054] There are property rights words such as professional terms, device structure design, legal terms and inventor's self-made phrases in the property registration text. When using the word segmentation algorithm to segment the sentences in step S1, the long words that meet the rights description may be segmented into multiple short phrases, and the property rights words cannot be accurately segmented. Therefore, the possibility of different word segments constituting property rights words is measured by the number of occurrences of different word segments in each sentence in the other sentences, and then the different word segments are merged to obtain the characteristic words of the analysis text.

[0055] Preferably, in some possible implementation modes of the embodiments of the present invention, the method for acquiring characteristic vocabulary includes: arranging the segmentation words of each sentence in the analysis text in order of position to obtain a segmentation sequence of the corresponding sentence; for the segmentation sequence of each sentence in the analysis text, randomly selecting a segmentation word in the segmentation sequence as a target segmentation word, and the target segmentation word and each subsequent segmentation word constitute a segmentation pair of the target segmentation word; in the analysis text, the proportion of the number of sentences of each segmentation pair containing the target segmentation word in the number of sentences containing the target segmentation word is used as the characteristic degree of each segmentation pair of the target segmentation word; selecting the segmentation pair corresponding to the maximum value in the characteristic degree of the target segmentation word, and recording it as the initial characteristic vocabulary of the target segmentation word; recording the initial characteristic vocabulary of all segmentations of the sentences in the analysis text as the characteristic vocabulary in the analysis text.

[0056] If the frequency of the two segmented words in each segmented word pair of the target segmented word in the same combination in the analyzed text is higher, the possibility that the two segmented words in the analyzed pair form a long word that is a property right word in the property right registration text is greater, and the characteristic degree of the segmented word pair is greater. In order to reduce the amount of calculation and ensure the accuracy of the property right word extraction in the property right registration text, the segmented word pair corresponding to the maximum value in the characteristic degree of the target segmented word is selected as the property right word formed based on the target segmented word, and the characteristic vocabulary of the analyzed text is obtained.

[0057] It should be noted that the sentence containing each participle pair of the target participle means that both participles in the participle pair exist in all the participles of the sentence, and the first participle is located before the second participle; the sentence containing the target participle means that the target participle exists in all the participles of the sentence. The method for obtaining the initial feature vocabulary of all the participles in all the sentences in the analysis text is the same as the method for obtaining the initial feature vocabulary of the target sentence.

[0058] Step S3: Determine the recurring paragraphs of each characteristic word in the test text in the property text; obtain the overall subject relevance of each characteristic word in the test text to the property text based on the number of characters and the number of words in each recurring paragraph of the property text where the characteristic words in the test text appear, and the similarity between the entity relationships between the clauses in which the characteristic words of the test text appearing in the recurring paragraphs are located and the remaining clauses in the property text.

[0059] In order to analyze the infringement of the property text by the test text, considering the similarity between the test text and the property text, it is necessary to determine the paragraphs in the property text where the characteristic words of the test text appear, and obtain the recurring paragraphs of the characteristic words in the property text. In the embodiment of the present invention, the recurring paragraphs contain two segment words in the segmentation pair corresponding to each characteristic word of the test text.

[0060] The similarity of the entity relationship between the sentence where the characteristic words are located and the rest of the sentences in the analysis text reflects the importance of the characteristic words in the analysis text. The number of characters and words in which the characteristic words in the test text appear in the recurring paragraphs of each characteristic word in the property text shows the possibility that the main content of the test text appears in the recurring paragraph of the characteristic words; the similarity of the entity relationship between the sentence where the characteristic words of the test text appear in the recurring paragraph and the rest of the sentences in the property text shows the possibility that the main content of the test text is the main content of the property text; the two factors are combined to judge the degree of correlation between the characteristic words and the main content of the property text, and the overall main correlation is obtained.

[0061] See also Figure 2 , which shows a flowchart of a method for obtaining the overall subject relevance provided by an embodiment of the present invention, the method comprising:

[0062] Step S310: Obtain the lexical importance index of each characteristic word in the analysis text according to the similarity between the entity relationships between the clause where each characteristic word is located and the other clauses in the analysis text, and the distribution position of the other clauses in the analysis text.

[0063] Different parts of the property registration text have different description focuses. Under normal circumstances, important descriptive words in the property registration text not only frequently appear in the entire text with similar entity relationships, but also are usually located in the front part of the property registration text. Based on the above two features, the lexical importance index of the feature words is obtained.

[0064] Preferably, in some possible implementations of the embodiments of the present invention, the method for obtaining the important vocabulary index includes: extracting entity relations from each sentence of the analyzed text to obtain the SPO triple of the corresponding sentence; recording the proportion of the number of characters of the SPO triple in the total number of characters of its corresponding sentence as the sentence ratio of the SPO triple; optionally recording a characteristic vocabulary in the analyzed text as an example vocabulary, selecting similar triples of the example vocabulary from the SPO triples of each sentence in the analyzed text except the sentence where the example vocabulary is located, and comparing the similar triples with any SPO triple of the sentence where the example vocabulary is located. The first element is equal and the third element is equal; the mean of the sentence ratios of all similar triplets of the example vocabulary is used as the entity similarity index of the example vocabulary; the sentences in the analysis text are numbered, and the mean of the sequence numbers of the sentences in which all similar triplets of the example vocabulary are located is calculated as the position distribution index of the example vocabulary; the lexical importance index of the example word is obtained according to the number of similar triplets of the example vocabulary, the entity similarity index and the position distribution index; the number of similar triplets of the example vocabulary and the entity similarity index are both positively correlated with the lexical importance index, and the position distribution index is negatively correlated with the lexical importance index.

[0065] The SPO triples of a clause reflect the relationship between the entities of the clause. A single sentence usually corresponds to one SPO triple, and a complex context clause may correspond to multiple SPO triples. Because the relationship between important rights description words and entities is usually located in the subject and object of the clause in the narrative process, in order to analyze the similarity of the relationship between entities in different clauses, the similar words of the example vocabulary are determined based on the first element and the third element in different SPO triples. The larger the sentence ratio of the similar triples of the example vocabulary, the more similar the relationship between entities between the clause where the example vocabulary is located and the other clauses. At the same time, the more similar triples of the example vocabulary are, it means that the example vocabulary frequently appears in the analyzed text with similar relationship between entities. In this case, the possibility that the example vocabulary is an important rights description vocabulary is greater, and the vocabulary importance index is greater. If the position distribution index of the example vocabulary is smaller, it means that the position of the example vocabulary in the analyzed text is closer, and the possibility that the example vocabulary is an important rights description vocabulary is greater, and the vocabulary importance index is greater.

[0066] In a specific implementation of the embodiment of the present invention, assuming that the example segmentation is the ath characteristic word in the analyzed text, the word importance index is expressed by the formula:

[0067]

[0068] In the formula, CZ_F a SX is the lexical importance index for analyzing the ath feature word in the text; a is the entity similarity index of the ath feature word in the analysis text; N ais the number of similar triples of the ath feature word in the analysis text; W a It is an indicator for analyzing the position distribution of the ath feature word in the text.

[0069] According to the above method, the vocabulary importance indexes of all characteristic words of the analyzed text are obtained.

[0070] In a specific implementation of the embodiment of the present invention, an entity relationship extraction algorithm is selected to obtain the SPO triples of the clauses.

[0071] It should be noted that the method for obtaining the lexical importance index of all characteristic words in the analysis text is the same as the method for obtaining the lexical importance index of the example word segmentation. The number of characters in the SPO triple is equal to the sum of the number of characters of the three elements in the triple. The sequence numbers of the clauses in the analysis text increase from sequence number 1, and the closer the clause is to the beginning of the analysis text, the smaller the sequence number.

[0072] Step S320: Obtain the overall subject relevance of each characteristic word in the test text to the property text based on the number of characters and the number of words in each characteristic word in the test text that appear in the recurring paragraph of the property text, and the lexical importance index of the characteristic words in the test text that appear in the recurring paragraph.

[0073] The same words may appear in different property registration texts, but the important contents may be quite different. For example, the focus of the first property registration text is to improve a certain technical means, and the second property registration text only quotes the technical means to achieve its own important contents. In this case, the same words may appear in the two property registration texts, but the first property registration text may contain a large number of key words and important technical means words in the same field. The important technical means words in the second property registration text have a lower correlation with the important contents of the text, and the important technical means words usually appear in local paragraphs of the second property registration text.

[0074] Therefore, assuming that the vth characteristic word of the test text appears in the bth paragraph of the property text, if the main content of the test text appears less in the bth paragraph and the bth paragraph is not the main content of the property text, that is, the fewer the number of characters and the number of words of the characteristic words in the test text in the bth paragraph, the lower the importance of the characteristic words in the test text appearing in the bth paragraph in the property text, that is, the smaller the lexical importance index, it means that the vth characteristic word is more likely to be called by the property text, and the lower the correlation between the vth characteristic word and the main content of the property text.

[0075] Preferably, in some possible implementations of the embodiments of the present invention, the method for obtaining the overall subject relevance includes: randomly selecting a characteristic word in the test text as the target word, taking the ratio of the total number of characters of the characteristic words in the test text that appear in each recurring paragraph of the target word in the property text to the total number of characters in the recurring paragraph as the character repetition of the target word in each recurring paragraph of the property text; selecting characteristic words in the test text that appear in each recurring paragraph of the target word in the property text as the focus words of the recurring paragraph; taking the ratio of the mean value of the lexical importance index of the focus words of the target word in each recurring paragraph of the property text to the maximum value of the lexical importance index of the characteristic words in the property text as the text importance of the target word in each recurring paragraph of the property text; obtaining the local subject relevance of the target word in each recurring paragraph of the property text according to the character repetition, the text importance, and the number of words of the characteristic words in the test text that appear in each recurring paragraph of the target word in the property text; taking the mean value of the local subject relevance of the target word in all recurring paragraphs of the property text as the overall subject relevance of the target word relative to the property text. Assuming that the target word is the vth characteristic word in the test text, the overall subject relevance is expressed by the formula:

[0076]

[0077] In the formula, GZ v is the overall subject relevance of the vth characteristic word in the test text relative to the property text; U v GJ is the total number of recurring paragraphs of the vth characteristic word in the test text in the property text; v,u is the local subject relevance of the vth characteristic word in the test text in the uth recurring paragraph of the property text; P v,u CZ_C is the character repetition degree of the vth characteristic word in the test text in the uth recurring paragraph of the property text; v,u CZ_C is the mean of the lexical importance indexes of all the focus words in the u-th recurring paragraph of the property text for the v-th feature word in the test text; max is the maximum value of the vocabulary importance index of the characteristic vocabulary in the property rights text; v,u is the number of words that appear in the u-th recurring paragraph of the property rights text among the feature words in the test text; exp is an exponential function with a natural constant as the base.

[0078] It should be noted that P v,u With Num v,u The smaller the value, the less the main content of the text to be tested appears in the u-th recurrence paragraph. The smaller the value, the higher the probability that the vth feature word is called by the property text, which makes the correlation between the vth feature word and the main content of the property text lower. The overall subject correlation GZ v The smaller it is, the lower the reliability of the v-th feature word in identifying infringement of property rights text. v,u The lexical importance index of the focus words in the u-th recurring paragraph in the meaning is obtained by analyzing the similarity of the entity relationships between the characteristic words in the property text.

[0079] Step S4: Based on the overall subject correlation difference and text difference between the characteristic vocabulary of the text to be tested and the property text, obtain the infringement index of the text to be tested on the property text; based on the infringement index, protect the data integrity of the blockchain.

[0080] In order to maintain the integrity of the property registration text data, it is necessary to detect infringement of key information in the test text, that is, the higher the repetition of the main information in the test text and the property text, the greater the possibility that the test text infringes the property text. Through the overall subject correlation difference and text difference of the characteristic vocabulary of the test text and the property text, the information similarity between the test text and the property text is analyzed to obtain the infringement index of the test text on the property text.

[0081] See also Figure 3 , which shows a flowchart of a method for obtaining infringement indicators provided by an embodiment of the present invention, the method comprising:

[0082] Step S410: The ratio of the frequency of occurrence of each characteristic word of the text to be tested in the text to be tested to the average frequency of occurrence of the characteristic words in all property texts on the blockchain is used as the block association index of each characteristic word of the text to be tested; negative correlation mapping is performed on the block association index of each characteristic word of the text to be tested, and the product of the mapping result and the overall subject association is used as the recognition support index of each characteristic word of the text to be tested.

[0083] In a specific implementation of the embodiment of the present invention, the identification support index is expressed by the formula:

[0084]

[0085] In the formula, SB v is the recognition support index of the vth feature word in the test text; L v is the frequency of occurrence of the vth characteristic word in the text to be tested; is the average of the occurrence frequencies of the vth characteristic word in the test text in all property rights texts on the blockchain; GZ is the block association index of the vth feature word in the test text; v is the overall subject relevance of the vth characteristic word in the test text relative to the property rights text.

[0086] It should be noted that if The larger the value, the greater the possibility that the vth feature word in the test text is an important content in the test text but not in the property text on the blockchain. The smaller the correlation between the vth feature word and the property text on the blockchain, the greater the possibility that the vth feature word is called by all property texts. At the same time, the overall subject correlation GZ v The larger the value, the greater the possibility that the vth feature word is called by all property texts, and the lower the reliability of the vth feature word for infringement identification of property texts. v The smaller.

[0087] Step S420: perform hash processing on the test text to obtain the hash value of the test text; obtain the difference sequence number, the hash values ​​of the test text and the property text are different in the characters of each difference sequence number; perform word vector conversion on the feature words corresponding to each difference sequence number of the hash value of the analysis text to obtain the feature vector of the hash value of the analysis text at each difference sequence number; take the average of the distances between the feature vectors of the hash values ​​of the test text and the property text at the same difference sequence number as the semantic difference value between the test text and the analysis text.

[0088] The hash value of the text to be tested is obtained based on the principle of perceptual hash function. The specific acquisition method includes: arranging the recognition support index of the characteristic vocabulary of the text to be tested in the order of the characteristic vocabulary position to obtain the recognition sequence of the text to be tested; performing one-dimensional DCT transformation on the recognition sequence to obtain the DCT coefficient of each element in the recognition sequence; calculating the mean of the DCT coefficients of all elements in the recognition sequence as the classification coefficient; setting the encoding of the characteristic vocabulary corresponding to the element whose DCT coefficient in the recognition sequence is greater than the classification coefficient to 1, and setting the encoding of the characteristic vocabulary corresponding to the element whose DCT coefficient is less than or equal to the classification coefficient to 0; arranging the encoding of the characteristic vocabulary of the text to be tested in the order of the characteristic vocabulary position to form a character string as the hash value of the text to be tested. Among them, performing one-dimensional DCT transformation on the sequence is a well-known technology and will not be repeated here.

[0089] The hash value is equivalent to a string consisting of a series of 0s and 1s. Each characteristic word in the analysis text corresponds to a DCT coefficient, and each DCT coefficient corresponds to a character in the hash value. Then, each character in the hash value corresponds to a characteristic word in the analysis text. If the difference in the characteristic words corresponding to the characters of each difference sequence number in the hash value of the analysis text is greater, it means that the semantic similarity between the test text and the characteristic words in the analysis text is smaller, and the semantic difference value is greater.

[0090] In a specific implementation of the embodiment of the present invention, the word2vec model is used to convert the feature vocabulary into a word vector to obtain a feature vector.

[0091] It should be noted that the method for obtaining the hash values ​​of the test text and the property text is the same. In the process of obtaining the hash value of the test text, the test text is replaced by the property text, and the property text is replaced by the test text. The distance between two feature vectors refers to the Euclidean distance.

[0092] Step S430: taking the ratio of the number of difference serial numbers between the test text and the property text to the target number of serial numbers as the hash difference value between the test text and the property text; obtaining the infringement index of the test text on the property text according to the hash difference value and the semantic difference value.

[0093] In a specific implementation of the embodiment of the present invention, the infringement index is expressed by the formula:

[0094]

[0095] In the formula, Qin is the infringement index of the test text and the property text; D is the semantic difference value between the test text and the property text; L is the number of difference serial numbers between the test text and the property text; L_m is the number of target serial numbers; is the hash difference between the text to be tested and the property text; exp is an exponential function with a natural constant as the base. It should be noted that if The larger the value of D is, the greater the difference in hash value between the test text and the property text and the greater the difference in semantic content, which means that the text similarity between the test text and the property text is smaller, and the possibility that the test text infringes the property text is smaller, and the infringement index is smaller.

[0096] In a specific implementation of the embodiment of the present invention, the number of target serial numbers is equal to the minimum value of the total number of characters of the hash values ​​of the text to be tested and the property rights text.

[0097] The blockchain may store multiple property rights texts. The property rights texts described before the current position represent a single property rights text on the blockchain. The infringement indicators of the text to be tested on all property rights texts on the blockchain are obtained in the same way as the infringement indicators of the text to be tested on the property rights text.

[0098] Based on the infringement index, determine whether the text to be tested infringes the property rights text on the blockchain, and store the text to be tested that does not infringe on the blockchain, thereby ensuring the integrity of the property rights registration text data stored on the blockchain. The specific operations are:

[0099] Determine whether the infringement indicators of the text to be tested for all property texts are less than the preset infringement threshold. If so, the text to be tested will not infringe the property text on the blockchain, and the text to be tested will be stored on the node of the blockchain; otherwise, the text to be tested cannot be stored on the node of the blockchain.

[0100] In a specific implementation of the embodiment of the present invention, the preset infringement threshold is set to 0.5.

[0101] So far, the present invention is completed.

[0102] Embodiment 2:

[0103] The present invention also proposes a computer device schematic diagram of a data integrity protection device based on hash function and blockchain technology, please refer to Figure 4 The computer device includes a memory 501, a processor 502, and a computer program 503 stored in the memory 501 and running on the processor 502, wherein when the processor 502 executes the computer program 503, the computer device can execute any of the data integrity protection methods based on hash functions and blockchain technology introduced above.

[0104] In addition, an embodiment of the present application also protects a device, which may include a memory and a processor, wherein the memory stores an executable program code, and the processor is used to call and execute the executable program code to execute a data integrity protection method based on hash function and blockchain technology provided in an embodiment of the present application.

[0105] In this embodiment, the functional modules of the device can be divided according to the above method example. For example, each functional module can be corresponded, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is schematic and is only a logical function division. There may be other division methods in actual implementation.

[0106] In the case of dividing each module according to each function, the device may also include a communication module, a signal analysis module, a complexity analysis module, a positioning module, etc. It should be noted that all relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module, which will not be repeated here.

[0107] It should be understood that the device provided in this embodiment is used to execute the above-mentioned data integrity protection method based on hash function and blockchain technology, and therefore can achieve the same effect as the above-mentioned implementation method.

[0108] In the case of an integrated unit, the device may include a processing module and a storage module. When the device is applied to a device, the processing module may be used to control and manage the actions of the device. The storage module may be used to support the device to execute mutual program codes, etc.

[0109] The processing module may be a processor or a controller, which may implement or execute various exemplary logic blocks, modules and circuits included in the disclosure of the present application. The processor may also be a combination that implements a computing function, such as a combination of one or more microprocessors, a combination of a digital signal processing (DSP) and a microprocessor, etc. The storage module may be a memory.

[0110] Embodiment 3:

[0111] This embodiment also provides a computer-readable storage medium, which stores computer program code. When the computer program code runs on a computer, the computer executes the above-mentioned related method steps to implement a data integrity protection method based on hash function and blockchain technology provided in the above embodiment.

[0112] Embodiment 4:

[0113] This embodiment also provides a computer program product. When the computer program product runs on a computer, it enables the computer to execute the above-mentioned related steps to implement a data integrity protection method based on hash function and blockchain technology provided in the above embodiment.

[0114] Among them, the device, computer-readable storage medium, computer program product or chip provided in this embodiment is used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be repeated here.

[0115] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0116] It should be noted that the sequence of the above embodiments of the present invention is only for description and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0117] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.

Claims

1. A data integrity protection method based on hash function and blockchain technology, characterized in that: The method includes: Obtain the text to be tested and the property rights text stored on the blockchain, and record the text to be tested and the property rights text as analysis text; divide the analysis text into different sentences, and segment the sentences to obtain segmented words; According to the number of occurrences of each participle of each sentence in the analyzed text and the other participles in the other sentences, characteristic words in the analyzed text are obtained; Determine the recurring paragraphs of each characteristic word in the test text in the property text; obtain the overall subject relevance of each characteristic word in the test text to the property text based on the number of characters and the number of words in each recurring paragraph of the characteristic word in the test text, and the similarity between the entity relationship between the sentence in which the characteristic word in the test text appears in the recurring paragraph and the other sentences in the property text; According to the overall subject correlation difference and text difference between the characteristic vocabulary of the text to be tested and the property text, the infringement index of the text to be tested on the property text is obtained; based on the infringement index, the data integrity of the blockchain is protected.

2. The data integrity protection method based on hash function and blockchain technology according to claim 1 is characterized in that: The acquiring and analyzing characteristic words in the text includes: Arrange the participles of each sentence in the analyzed text in order of position to obtain the participle sequence of the corresponding sentence; For each sentence segmentation sequence of the analyzed text, a segmentation in the segmentation sequence is randomly selected and recorded as a target segmentation, and the target segmentation and each segmentation after it constitute a segmentation pair of the target segmentation; in the analyzed text, the proportion of the number of sentences of each segmentation pair containing the target segmentation in the number of sentences containing the target segmentation is used as the characteristic degree of each segmentation pair of the target segmentation; the segmentation pair corresponding to the maximum value of the characteristic degree of the target segmentation is selected and recorded as the initial characteristic vocabulary of the target segmentation; The initial characteristic words of all the word segments in the analyzed text are recorded as the characteristic words in the analyzed text.

3. The data integrity protection method based on hash function and blockchain technology according to claim 1 is characterized in that: The step of obtaining the overall subject relevance of each characteristic word of the text to be tested relative to the property text includes: According to the similarity of the entity relationship between the clause where each characteristic word is located and the other clauses in the analyzed text, and the distribution position of the other clauses in the analyzed text, the lexical importance index of each characteristic word in the analyzed text is obtained; The overall subject relevance of each characteristic word in the test text to the property text is obtained based on the number of characters and the number of words in each characteristic word's recurring paragraph in the property text, as well as the lexical importance index of the characteristic words in the test text that appear in the recurring paragraph.

4. The data integrity protection method based on hash function and blockchain technology according to claim 3 is characterized in that: The acquisition of the important vocabulary index of each characteristic vocabulary in the analysis text includes: Entity relations are extracted for each sentence of the analyzed text to obtain the SPO triples of the corresponding sentences; the proportion of the number of characters of the SPO triples in the total number of characters of the corresponding sentences is recorded as the sentence ratio of the SPO triples; A characteristic word in the analysis text is recorded as an example word, and a similar triplet of the example word is selected from the SPO triples of each sentence in the analysis text except the sentence where the example word is located, and the similar triplet is equal to the first element and the third element of any SPO triplet of the sentence where the example word is located; the average of the sentence ratios of all similar triples of the example word is used as the entity similarity index of the example word; Number the sentences in the analyzed text, and calculate the mean of the sequence numbers of the sentences where all similar triples of the sample words are located, as the position distribution indicator of the sample words; According to the number of similar triples of the example vocabulary, the entity similarity index and the position distribution index, the vocabulary importance index of the example word segmentation is obtained; the number of similar triples of the example vocabulary and the entity similarity index are both positively correlated with the vocabulary importance index, and the position distribution index is negatively correlated with the vocabulary importance index.

5. The data integrity protection method based on hash function and blockchain technology according to claim 3 is characterized in that: The step of obtaining the overall subject relevance of each characteristic word in the test text to the property text includes: A characteristic word in the test text is randomly selected as the target word, and the ratio of the total number of characters of the characteristic word in the test text that appears in each recurring paragraph of the target word in the property text to the total number of characters in the recurring paragraph is taken as the character repetition degree of the target word in each recurring paragraph of the property text; Select characteristic words that appear in each recurring paragraph of the target word in the property text in the test text, and record them as the focus words of the recurring paragraph; take the ratio of the mean value of the lexical importance index of the focus words in each recurring paragraph of the target word in the property text to the maximum value of the lexical importance index of the characteristic words in the property text as the text importance of the target word in each recurring paragraph of the property text; According to the character repetition degree, the text importance, and the number of words in which the characteristic words in the test text appear in each recurring paragraph of the target word in the property text, the local subject relevance of the target word in each recurring paragraph of the property text is obtained; The average of the local subject relevance of the target vocabulary in all recurring paragraphs of the property text is used as the overall subject relevance of the target vocabulary relative to the property text.

6. The data integrity protection method based on hash function and blockchain technology according to claim 1 is characterized in that: The step of obtaining the infringement index of the tested text on the property text includes: The ratio of the frequency of occurrence of each characteristic word of the test text in the test text to the average frequency of occurrence of the characteristic word in all property texts on the blockchain is used as the block association index of each characteristic word of the test text; the block association index of each characteristic word of the test text is negatively correlated, and the product of the mapping result and the overall subject association degree is used as the recognition support index of each characteristic word of the test text; Perform hash processing on the text to be tested to obtain the hash value of the text to be tested; obtain the difference sequence number, the hash values ​​of the text to be tested and the property text are different in the characters of each difference sequence number; perform word vector conversion on the feature words corresponding to the characters of the hash value of the analysis text at each difference sequence number to obtain the feature vector of the hash value of the analysis text at each difference sequence number; take the average value of the distance between the feature vectors of the hash values ​​of the text to be tested and the property text at the same difference sequence number as the semantic difference value between the text to be tested and the analysis text; The ratio of the number of difference serial numbers between the text to be tested and the property text to the number of target serial numbers is used as the hash difference value between the text to be tested and the property text; based on the hash difference value and the semantic difference value, the infringement index of the text to be tested on the property text is obtained.

7. The data integrity protection method based on hash function and blockchain technology according to claim 6 is characterized in that: The step of performing hash processing on the text to be tested to obtain a hash value of the text to be tested includes: Arranging the recognition support indicators of the characteristic words of the text to be tested in order of the characteristic words' positions to obtain a recognition sequence of the text to be tested; performing a one-dimensional DCT transformation on the recognition sequence to obtain a DCT coefficient of each element in the recognition sequence; Calculate the mean of the DCT coefficients of all elements in the recognition sequence as the classification coefficient; set the encoding of the feature vocabulary corresponding to the elements in the recognition sequence whose DCT coefficient is greater than the classification coefficient to 1, and set the encoding of the feature vocabulary corresponding to the elements whose DCT coefficient is less than or equal to the classification coefficient to 0; arrange the encoding of the feature vocabulary of the text to be tested in the order of the position of the feature vocabulary to form a string as the hash value of the text to be tested.

8. The data integrity protection method based on hash function and blockchain technology according to claim 1 is characterized in that: Based on the infringement indicator, protecting the data integrity of the blockchain includes: Determine whether the infringement indicators of the text to be tested for all property texts are less than the preset infringement threshold. If so, store the text to be tested on the node of the blockchain. Otherwise, the text to be tested cannot be stored on the node of the blockchain.

9. The data integrity protection method based on hash function and blockchain technology according to claim 2 is characterized in that: Each characteristic word in the text to be tested contained in the recurring paragraph corresponds to two participles in the participle pair.

10. The data integrity protection method based on hash function and blockchain technology according to claim 6 is characterized in that: The number of target serial numbers is equal to the minimum value of the total number of characters in the hash values ​​of the text to be tested and the property rights text.

Citation Information

Patent Citations

  • Electronic case duplicate checking method and device based on word segmentation text and computer equipment

    CN111814447A

  • Method and system for efficiently extracting key data information of archives

    CN117973387A

  • Scientific and technological achievement authenticity verification and evaluation platform based on block chain

    CN119149503A

  • Data processing methods, apparatuses, and devices

    US20210326357A1