Data integrity protection method based on hash function and blockchain technology
By combining hash functions with blockchain technology, the correlation and difference of characteristic words between the text to be tested and the property text are obtained, and the infringement index is calculated, which solves the problem of inaccurate infringement judgment in existing technologies and achieves more efficient data integrity protection.
Patent Information
- Application Number
- CN202510178144.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-02-18
AI Technical Summary
When existing methods judge infringement of property rights data, the word frequency-inverse document frequency algorithm does not highlight key words, resulting in low accuracy of infringement judgment.
A method based on hash functions and blockchain technology is used to obtain the characteristic vocabulary of the text to be tested and the property text, analyze their overall subject correlation and text differences, calculate the infringement index, and determine whether the text to be tested is stored on the blockchain.
It improves the accuracy of property rights data infringement judgment, avoids malicious modification of the text to be tested before storage, and protects the integrity of blockchain data.
Smart Images

Figure CN120105489B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of protecting data integrity, and in particular to a data integrity protection method based on hash functions and blockchain technology. Background Art
[0002] When protecting data in the property registration system, blockchain technology is usually used to merge data blocks to form an encrypted chain data structure. The characteristics of decentralized trust are used to improve the traditional method of single point failure or excessive load on the central node, which leads to data loss, damage and leakage in the system. At the same time, the blockchain has the characteristic of being tamper-proof, so the application of blockchain technology in the property registration system can significantly improve the security and reliability of property data.
[0003] Before property rights data is stored on the blockchain, it could be maliciously copied or modified, potentially infringing the user's property rights stored there. Existing methods typically use a term frequency-inverse document frequency algorithm to determine whether property rights data infringes. However, when this algorithm screens key terms, some terms that represent property rights validity are not prominent, reducing the accuracy of infringement judgments. Summary of the Invention
[0004] In order to solve the technical problem that the words representing the validity of property rights are not prominent in property rights data, resulting in low accuracy in infringement judgment of property rights data, the purpose of the present invention is to provide a data integrity protection method based on hash functions and blockchain technology. The technical solution adopted is as follows:
[0005] The present invention proposes a data integrity protection method based on hash function and blockchain technology, the method comprising:
[0006] Obtain the text to be tested and the property rights text stored on the blockchain, and record the text to be tested and the property rights text as analysis text; divide the analysis text into different sentences, and segment the sentences to obtain segmented words;
[0007] Obtain characteristic words in the analyzed text according to the number of occurrences of each participle in each sentence of the analyzed text and the other participles in the other sentences;
[0008] Determine the paragraphs in the property text in which each characteristic word in the test text recurs; obtain the overall subject relevance of each characteristic word in the test text relative to the property text based on the number of characters and the number of words in each paragraph in which the characteristic word appears, and the similarity between the entity relationships between the clause in which the characteristic word in the test text appears in the recurring paragraph and the remaining clauses in the property text;
[0009] Based on the overall subject correlation difference and text difference of the characteristic vocabulary of the text to be tested and the property text, an infringement index of the text to be tested on the property text is obtained; based on the infringement index, the data integrity of the blockchain is protected.
[0010] Furthermore, the acquisition of characteristic words in the analysis text includes:
[0011] Arrange the segmented words of each sentence in the analyzed text in order of position to obtain the segmented word sequence of the corresponding sentence;
[0012] For each sentence segmentation sequence of the analyzed text, a segmentation in the segmentation sequence is randomly selected and recorded as the target segmentation, and the target segmentation and each subsequent segmentation form a segmentation pair of the target segmentation; in the analyzed text, the proportion of the number of sentences containing each segmentation pair of the target segmentation to the number of sentences containing the target segmentation is used as the characteristic degree of each segmentation pair of the target segmentation; the segmentation pair corresponding to the maximum value of the characteristic degree of the target segmentation is selected and recorded as the initial characteristic vocabulary of the target segmentation;
[0013] The initial characteristic words of all the segmented words in the analyzed text are recorded as the characteristic words in the analyzed text.
[0014] Furthermore, obtaining the relevance of each characteristic word of the text to be tested to the overall subject of the property rights text includes:
[0015] Obtaining the lexical importance index of each characteristic word in the analysis text based on the similarity of the entity relationship between the clause where each characteristic word is located and the remaining clauses in the analysis text, as well as the distribution position of the remaining clauses in the analysis text;
[0016] Based on the number of characters and the number of words in the repeated paragraphs of the property text in which the characteristic words in the text to be tested appear, and the lexical importance index of the characteristic words in the text to be tested that appear in the repeated paragraphs, the overall main body relevance of each characteristic word in the text to be tested relative to the property text is obtained.
[0017] Furthermore, the acquisition of the vocabulary importance index of each characteristic word in the analysis text includes:
[0018] Entity relations are extracted for each sentence in the analyzed text to obtain the SPO triples of the corresponding sentences. The ratio of the number of characters in the SPO triple to the total number of characters in the corresponding sentence is recorded as the sentence ratio of the SPO triple.
[0019] A characteristic word in the analysis text is randomly recorded as an example word, and similar triples of the example word are selected from the SPO triples of each sentence in the analysis text except the sentence where the example word is located. The similar triples have the same first element and the same third element as any SPO triple of the sentence where the example word is located; the average of the sentence ratios of all similar triples of the example word is used as the entity similarity index of the example word;
[0020] Number the sentences in the analyzed text and calculate the mean of the sequence numbers of the sentences where all similar triples of the example words are located as the position distribution indicator of the example words;
[0021] The lexical importance index of the example word segmentation is obtained based on the number of similar triples of the example vocabulary, the entity similarity index and the position distribution index; the number of similar triples of the example vocabulary and the entity similarity index are both positively correlated with the lexical importance index, and the position distribution index is negatively correlated with the lexical importance index.
[0022] Furthermore, the step of obtaining the overall subject relevance of each characteristic word in the test text to the property text includes:
[0023] A characteristic word in the test text is randomly selected as the target word, and the ratio of the total number of characters in which the characteristic word in the test text appears in each recurring paragraph of the target word in the property text to the total number of characters in the recurring paragraph is used as the character repetition rate of the target word in each recurring paragraph of the property text;
[0024] Selecting characteristic words that appear in each recurring paragraph of the target word in the property text in the test text as the focus words of the recurring paragraph; taking the ratio of the mean value of the lexical importance index of the focus words in each recurring paragraph of the target word in the property text to the maximum value of the lexical importance index of the characteristic words in the property text as the text importance of the target word in each recurring paragraph of the property text;
[0025] Obtaining the local subject relevance of the target word in each recurring paragraph of the property text based on the character repetition degree, the text importance, and the number of times the characteristic words in the test text appear in each recurring paragraph of the target word in the property text;
[0026] The average of the local subject relevance of the target vocabulary in all recurring paragraphs of the property text is used as the overall subject relevance of the target vocabulary relative to the property text.
[0027] Furthermore, obtaining the infringement index of the tested text on the property rights text includes:
[0028] The ratio of the frequency of occurrence of each characteristic word in the test text to the average frequency of occurrence of the characteristic word in all property texts on the blockchain is used as the block association index of each characteristic word in the test text; the block association index of each characteristic word in the test text is negatively correlated, and the product of the mapping result and the overall subject association degree is used as the recognition support index of each characteristic word in the test text;
[0029] Perform hash processing on the test text to obtain the hash value of the test text; obtain the difference sequence number, where the hash values of the test text and the property text are different at each character of the difference sequence number; perform word vector conversion on the feature words corresponding to the characters of each difference sequence number of the hash value of the analysis text to obtain the feature vector of the hash value of the analysis text at each difference sequence number; take the average value of the distance between the feature vectors of the hash values of the test text and the property text at the same difference sequence number as the semantic difference value between the test text and the analysis text;
[0030] The ratio of the number of difference serial numbers between the text to be tested and the property text to the target number of serial numbers is used as the hash difference value between the text to be tested and the property text; based on the hash difference value and the semantic difference value, the infringement index of the text to be tested on the property text is obtained.
[0031] Furthermore, performing hash processing on the text to be tested to obtain a hash value of the text to be tested includes:
[0032] Arranging the recognition support indicators of the characteristic words of the test text in order of their positions to obtain a recognition sequence of the test text; performing a one-dimensional DCT transform on the recognition sequence to obtain a DCT coefficient of each element in the recognition sequence;
[0033] Calculate the mean of the DCT coefficients of all elements in the recognition sequence as the classification coefficient; set the encoding of the feature vocabulary corresponding to the elements in the recognition sequence whose DCT coefficient is greater than the classification coefficient to 1, and set the encoding of the feature vocabulary corresponding to the elements whose DCT coefficient is less than or equal to the classification coefficient to 0; arrange the encoding of the feature vocabulary of the text to be tested in the order of the position of the feature vocabulary to form a string as the hash value of the text to be tested.
[0034] Furthermore, protecting the data integrity of the blockchain based on the infringement indicator includes:
[0035] Determine whether the infringement indicators of the text to be tested for all property rights texts are all less than the preset infringement threshold. If so, store the text to be tested on the node of the blockchain. Otherwise, the text to be tested cannot be stored on the node of the blockchain.
[0036] Furthermore, each characteristic word in the text to be tested contained in the recurring paragraph corresponds to two segmented words in the segmented word pair.
[0037] Furthermore, the number of target serial numbers is equal to the minimum value of the total number of characters in the hash values of the text to be tested and the property rights text.
[0038] The present invention has the following beneficial effects:
[0039] In an embodiment of the present invention, there are professional terms and property rights words such as device structure design in the analysis text, and the property rights words cannot be accurately segmented when the analysis text is segmented. The possibility of different segmentations constituting property rights words is measured by the number of times different segmentations in each sentence appear in other sentences, and the characteristic vocabulary of the analysis text is obtained; the number of characters and the number of vocabulary in which the characteristic vocabulary in the test text appears in the recurring paragraph of each characteristic vocabulary in the property rights text presents the possibility of the main content of the test text appearing in the recurring paragraph of the characteristic vocabulary, and the relationship between the entity of the sentence in which the characteristic vocabulary of the test text appears in the recurring paragraph and the other sentences in the property rights text is compared. Similar situations are presented, presenting the possibility that the main content of the test text is the main content of the property text, and comprehensively judging the degree of correlation between the characteristic vocabulary and the main content of the property text by combining the two factors to obtain the overall subject correlation; through the overall subject correlation difference and text difference of the characteristic vocabulary of the test text and the property text, the information similarity between the test text and the property text is analyzed, and the possibility of the test text infringing the property text is analyzed to obtain the infringement index, thereby improving the accuracy of the infringement judgment of the document; based on the infringement index, decide whether to store the test text on the blockchain, so as to avoid malicious modification of the property information before the test text is put on the chain, thereby avoiding infringement of the user's property rights. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0041] Figure 1 A flowchart of a data integrity protection method based on hash functions and blockchain technology provided by one embodiment of the present invention;
[0042] Figure 2 A flowchart of a method for obtaining overall subject relevance provided by one embodiment of the present invention;
[0043] Figure 3 A flowchart of a method for obtaining infringement indicators provided by one embodiment of the present invention;
[0044] Figure 4A schematic diagram of a computer device providing a data integrity protection device based on hash functions and blockchain technology, provided in accordance with one embodiment of the present invention. DETAILED DESCRIPTION
[0045] To further illustrate the technical means and effectiveness of the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the data integrity protection method based on hash functions and blockchain technology, including its specific implementation, structure, features, and effectiveness. In the following description, different references to "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.
[0046] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0047] The following describes in detail the specific scheme of the data integrity protection method based on hash function and blockchain technology provided by the present invention with reference to the accompanying drawings.
[0048] Example 1:
[0049] This invention proposes a data integrity protection method based on hash function and blockchain technology, please refer to Figure 1 , which shows a flowchart of a data integrity protection method based on hash function and blockchain technology provided by one embodiment of the present invention, the method comprising:
[0050] Step S1: Obtain the text to be tested and the property rights text stored on the blockchain, and record the text to be tested and the property rights text as analysis text; divide the analysis text into different sentences, and segment the sentences to obtain segmented words.
[0051] Specifically, the property registration text to be stored on the blockchain is recorded as the test text, and the property registration text already stored on the blockchain node is recorded as the property text. Multiple property texts are stored on the blockchain. To analyze whether the test text infringes on the property text, the sentences in the analysis text are segmented.
[0052] In an embodiment of the present invention, the segmentation method is as follows: first, the analysis text is converted into a text document format; then, the sentences in the analysis text are divided based on punctuation marks to obtain different sentences, that is, the sentences between two adjacent punctuation marks are recorded as a sentence; finally, each sentence is segmented using the Jieba word segmentation algorithm to obtain multiple word segments. The Jieba word segmentation algorithm is well known to those skilled in the art and will not be described in detail here.
[0053] Step S2: Acquire characteristic words in the analysis text according to the number of occurrences of each segment word in each sentence in the analysis text and the number of occurrences of the remaining segment words in the remaining sentences.
[0054] Property registration documents contain proprietary terms, such as professional terminology, device structure designs, legal terminology, and inventor-created phrases. The word segmentation algorithm used to segment the sentences in step S1 may break long words that correspond to the rights description into multiple short phrases, making it impossible to accurately segment the proprietary terms. Therefore, the probability that different participles in each sentence constitute proprietary terms is measured by the number of times they appear in the remaining sentences. These different participles are then merged to obtain the characteristic vocabulary of the analyzed text.
[0055] Preferably, in some possible implementations of the embodiments of the present invention, the method for obtaining characteristic vocabulary includes: arranging the segmentation words of each sentence in the analysis text in order of position to obtain a segmentation sequence of the corresponding sentence; for the segmentation sequence of each sentence in the analysis text, any segmentation word in the segmentation sequence is recorded as the target segmentation word, and the target segmentation word and each segmentation word after it constitute a segmentation pair of the target segmentation word; in the analysis text, the proportion of the number of sentences of each segmentation pair containing the target segmentation word in the number of sentences containing the target segmentation word is used as the characteristic degree of each segmentation pair of the target segmentation word; selecting the segmentation pair corresponding to the maximum value in the characteristic degree of the target segmentation word, and recording it as the initial characteristic vocabulary of the target segmentation word; recording the initial characteristic vocabulary of all segmentations of the sentences in the analysis text as the characteristic vocabulary in the analysis text.
[0056] The more frequently the two segmented words in each segmented word pair of the target segmented word appear in the same combination in the analyzed text, the greater the likelihood that the two segmented words in the analyzed pair form a long word that represents a property rights term in the property rights registration text, and the greater the characteristic degree of the segmented word pair. To reduce computational complexity and ensure the accuracy of property rights term extraction from property rights registration texts, the segmented word pair corresponding to the maximum characteristic degree of the target segmented word is selected as the property rights term constructed based on the target segmented word, thereby obtaining the characteristic vocabulary of the analyzed text.
[0057] It should be noted that a sentence containing a target participle is one in which both participles in the pair exist, with the first participle preceding the second. A sentence containing a target participle is one in which the target participle exists. The method for obtaining the initial feature vocabulary for all participles in all sentences in the analyzed text is the same as that for the target sentence.
[0058] Step S3: Determine the recurring paragraphs of each characteristic word in the test text in the property text; obtain the overall subject relevance of each characteristic word in the test text to the property text based on the number of characters and the number of words in each recurring paragraph of the property text in which the characteristic word in the test text appears, and the similarity between the entity relationships between the clauses in which the characteristic word in the test text appears in the recurring paragraphs and the remaining clauses in the property text.
[0059] To analyze whether the test text infringes upon the proprietary text, and considering the similarity between the test text and the proprietary text, it is necessary to identify the paragraphs in the proprietary text in which the characteristic vocabulary of the test text appears, and obtain the recurring paragraphs of the characteristic vocabulary in the proprietary text. In this embodiment of the present invention, the recurring paragraphs contain two segmented words from the segmented word pairs corresponding to each characteristic vocabulary of the test text.
[0060] The similarity between the entity relationships between the clauses containing the characteristic words and the remaining clauses in the analyzed text reflects the importance of the characteristic words in the analyzed text. The number of characters and words in each characteristic word's recurring paragraph in the property text indicates the likelihood that the main content of the test text will appear in the recurring paragraph of the characteristic word. The similarity between the entity relationships between the clauses containing the characteristic words in the property text and the remaining clauses in the recurring paragraph indicates the likelihood that the main content of the test text is the main content of the property text. Combining these two factors, the degree of relevance between the characteristic words and the main content of the property text is determined to obtain the overall main content relevance.
[0061] See also Figure 2 , which shows a flowchart of a method for obtaining the overall subject relevance provided by one embodiment of the present invention, the method comprising:
[0062] Step S310: Obtain the lexical importance index of each characteristic word in the analysis text based on the similarity between the entity relationships between the clause where each characteristic word is located and the remaining clauses in the analysis text, as well as the distribution positions of the remaining clauses in the analysis text.
[0063] Different parts of the property registration text have different description focuses. Under normal circumstances, important descriptive words in the property registration text not only appear frequently in the entire text with similar inter-entity relationships, but also are usually located in the front part of the property registration text. Based on the above two features, the lexical importance index of the characteristic words is obtained.
[0064] Preferably, in some possible implementations of the present invention, the method for obtaining the important vocabulary index includes: extracting entity relations from each sentence of the analyzed text to obtain the SPO triple of the corresponding sentence; recording the proportion of the number of characters of the SPO triple in the total number of characters of its corresponding sentence as the sentence ratio of the SPO triple; optionally recording a characteristic word in the analyzed text as an example word, selecting similar triples of the example word from the SPO triples of each sentence in the analyzed text except the sentence where the example word is located, and comparing the similar triples with any SPO triple of the sentence where the example word is located. The first element is equal and the third element is equal; the mean of the sentence ratios of all similar triplets of the example vocabulary is used as the entity similarity index of the example vocabulary; the sentences in the analysis text are numbered, and the mean of the sequence numbers of the sentences in which all similar triplets of the example vocabulary are located is calculated as the position distribution index of the example vocabulary; the lexical importance index of the example word is obtained according to the number of similar triplets of the example vocabulary, the entity similarity index and the position distribution index; the number of similar triplets of the example vocabulary and the entity similarity index are both positively correlated with the lexical importance index, and the position distribution index is negatively correlated with the lexical importance index.
[0065] The SPO triples of a clause reflect the inter-entity relationships of the clause. A single sentence usually corresponds to one SPO triple, while a clause in a complex context may correspond to multiple SPO triples. Because the relationship between important rights description words and entities is usually located in the subject and object of the clause during the narrative process, in order to analyze the similarity of the inter-entity relationships of different clauses, the similar words of the example vocabulary are determined based on the first and third elements in different SPO triples. The larger the clause ratio of the similar triples of the example vocabulary, the more similar the inter-entity relationships between the clause in which the example vocabulary is located and the rest of the clauses. At the same time, the more similar triples of the example vocabulary are, the more frequent the example vocabulary appears in the analyzed text with similar inter-entity relationships. In this case, the more likely the example vocabulary is an important rights description word, and the larger the vocabulary importance index. If the position distribution index of the example vocabulary is smaller, it means that the example vocabulary is closer to the front in the analyzed text. In this case, the more likely the example vocabulary is an important rights description word, and the larger the vocabulary importance index.
[0066] In a specific implementation of the embodiment of the present invention, assuming that the example segmentation is the ath characteristic word in the analysis text, the vocabulary importance index is expressed by the formula:
[0067]
[0068] Where, To analyze the important lexical index of the a-th feature word in the text; To analyze the entity similarity index of the a-th feature word in the text; To analyze the number of similar triples of the a-th feature word in the text; It is an indicator of the position distribution of the ath feature word in the analysis text.
[0069] According to the above method, the vocabulary importance indexes of all characteristic words in the analyzed text are obtained.
[0070] In a specific implementation of the embodiment of the present invention, an entity relationship extraction algorithm is selected to obtain the SPO triples of the clauses.
[0071] It should be noted that the method for obtaining lexical importance indicators for all characteristic words in the analysis text is the same as the method for obtaining lexical importance indicators for the example word segmentation. The number of characters in an SPO triple is equal to the sum of the number of characters in the three elements of the triple. The sequence number of the clauses in the analysis text increases from sequence number 1, and the sequence number of the clauses closer to the beginning of the analysis text decreases.
[0072] Step S320: Obtain the overall subject relevance of each characteristic word in the test text to the property text based on the number of characters and the number of words in each characteristic word in the test text that appear in the recurring paragraph of the property text, as well as the lexical importance index of the characteristic words in the test text that appear in the recurring paragraph.
[0073] The same words may appear in different property registration texts, but the important contents may be quite different. For example, the focus of the first property registration text is to improve a certain technical means, and the second property registration text only quotes the technical means to achieve its own important content. In this case, the same words will appear in the two property registration texts, but the first property registration text may contain a large number of key words and important technical means words in the same field. The important technical means words in the second property registration text have a low correlation with the important content of the text, and the important technical means words usually appear in local paragraphs of the second property registration text.
[0074] Therefore, assuming that the vth characteristic word of the test text appears in the bth paragraph of the property text, if the main content of the test text appears less in the bth paragraph and the bth paragraph is not the main content of the property text, that is, the fewer the number of characters and words of the characteristic words in the test text in the bth paragraph, the lower the importance of the characteristic words in the test text appearing in the bth paragraph in the property text, that is, the smaller the vocabulary importance index, it means that the vth characteristic word is more likely to be called by the property text, and the lower the degree of correlation between the vth characteristic word and the main content of the property text.
[0075] Preferably, in some possible implementations of the embodiments of the present invention, the method for obtaining the overall subject relevance includes: randomly selecting a characteristic word in the test text as the target word, taking the ratio of the total number of characters of the characteristic word in the test text that appears in each recurring paragraph of the target word in the property text to the total number of characters in the recurring paragraph as the character repetition of the target word in each recurring paragraph of the property text; selecting a characteristic word in the test text that appears in each recurring paragraph of the target word in the property text as the focus word of the recurring paragraph; taking the ratio of the average value of the lexical importance index of the focus word in each recurring paragraph of the target word in the property text to the maximum value of the lexical importance index of the characteristic word in the property text as the text importance of the target word in each recurring paragraph of the property text; obtaining the local subject relevance of the target word in each recurring paragraph of the property text based on the character repetition, text importance, and the number of words of the characteristic word in the test text that appears in each recurring paragraph of the target word in the property text; and taking the average value of the local subject relevance of the target word in all recurring paragraphs of the property text as the overall subject relevance of the target word relative to the property text. Assuming that the target word is the vth feature word in the test text, the overall subject relevance is expressed as follows:
[0076]
[0077]
[0078] Where, is the overall subject relevance of the vth feature word in the test text relative to the property rights text; is the total number of paragraphs in which the vth characteristic word in the test text appears in the property text; is the local subject relevance of the vth feature word in the test text in the uth recurring paragraph of the property rights text; is the character repetition degree of the vth characteristic word in the test text in the uth recurring paragraph of the property rights text; is the mean of the lexical importance indexes of all the focus words in the u-th recurring paragraph of the property rights text for the v-th feature word in the test text; It is the maximum value among the lexical importance indicators of characteristic words in the property rights text; is the number of characteristic words in the test text that appear in the u-th recurring paragraph of the property rights text; exp is an exponential function with a natural constant as the base.
[0079] It should be noted that and The smaller the value, the less the main content of the text to be tested appears in the u-th recurrence paragraph. The smaller the value, the higher the probability that the vth feature word is called by the property text, which makes the correlation between the vth feature word and the main content of the property text lower. The smaller it is, the lower the reliability of the vth feature word in identifying infringement of property rights texts. The lexical importance index of the focus words in the u-th recurring paragraph in the meaning is obtained by analyzing the similarity of the entity relationships of the characteristic words in the property rights text.
[0080] Step S4: Based on the overall subject correlation difference and text difference of the characteristic vocabulary of the text to be tested and the property rights text, obtain the infringement index of the text to be tested on the property rights text; based on the infringement index, protect the data integrity of the blockchain.
[0081] To maintain the integrity of property registration text data, infringement detection is required for key information in the test text. Specifically, the greater the repetition of key information between the test text and the property text, the greater the likelihood that the test text infringes on the property text. By analyzing the overall subject relevance differences and textual differences between the characteristic vocabulary of the test text and the property text, we analyze the degree of information similarity between the test text and the property text, generating an infringement index for the test text against the property text.
[0082] See also Figure 3 , which shows a flowchart of a method for obtaining infringement indicators provided by an embodiment of the present invention, the method comprising:
[0083] Step S410: The ratio of the frequency of occurrence of each characteristic word in the text to be tested to the average frequency of occurrence of the characteristic word in all property texts on the blockchain is used as the block association index of each characteristic word in the text to be tested; a negative correlation mapping is performed on the block association index of each characteristic word in the text to be tested, and the product of the mapping result and the overall subject association degree is used as the recognition support index of each characteristic word in the text to be tested.
[0084] In a specific implementation of the embodiment of the present invention, the identification support index is expressed as follows:
[0085]
[0086] Where, is the recognition support index of the vth feature word in the test text; is the frequency of occurrence of the vth feature word in the test text; is the average of the occurrence frequencies of the vth characteristic word in the test text in all property rights texts on the blockchain; is the block association index of the vth feature word in the test text; is the overall subject relevance of the vth feature word in the test text relative to the property rights text.
[0087] It should be noted that if The larger the value, the greater the possibility that the vth feature word in the test text is an important content in the test text but not in the property rights text on the blockchain. The smaller the correlation between the vth feature word and the property rights text on the blockchain, the greater the possibility that the vth feature word is called by all property rights texts. At the same time, the overall subject correlation The larger the value, the greater the possibility that the vth feature word is called by all property texts, and the lower the reliability of the vth feature word in identifying infringement of property texts. The smaller.
[0088] Step S420: Perform hash processing on the text to be tested to obtain the hash value of the text to be tested; obtain the difference serial number, the hash value of the text to be tested and the property text are different in the characters of each difference serial number; perform word vector conversion on the feature vocabulary corresponding to each difference serial number of the hash value of the analysis text to obtain the feature vector of the hash value of the analysis text at each difference serial number; take the average value of the distance between the feature vectors of the hash values of the text to be tested and the property text at the same difference serial number as the semantic difference value between the text to be tested and the analysis text.
[0089] The hash value of the test text is obtained based on the principle of perceptual hash function. The specific acquisition method includes: arranging the recognition support indicators of the characteristic words of the test text in the order of the characteristic word positions to obtain the recognition sequence of the test text; performing a one-dimensional DCT transform on the recognition sequence to obtain the DCT coefficient of each element in the recognition sequence; calculating the mean of the DCT coefficients of all elements in the recognition sequence as the classification coefficient; setting the encoding of the characteristic words corresponding to the elements in the recognition sequence whose DCT coefficients are greater than the classification coefficient to 1, and setting the encoding of the characteristic words corresponding to the elements whose DCT coefficients are less than or equal to the classification coefficient to 0; and arranging the encoding of the characteristic words of the test text in the order of the characteristic word positions to form a string as the hash value of the test text. Among them, performing a one-dimensional DCT transform on the sequence is a well-known technology and will not be repeated here.
[0090] A hash value is equivalent to a string of 0s and 1s. Each characteristic word in the analysis text corresponds to a DCT coefficient, and each DCT coefficient corresponds to a character in the hash value. Therefore, each character in the hash value corresponds to a characteristic word in the analysis text. The greater the difference in the characteristic words corresponding to each character in the hash value of the analysis text, the less semantic similarity there is between the test text and the characteristic words in the analysis text, and the greater the semantic difference value.
[0091] In a specific implementation of the embodiment of the present invention, the word2vec model is used to convert the feature vocabulary into a word vector to obtain a feature vector.
[0092] It should be noted that the hash values of the test text and the property rights text are obtained in the same way. During the hash value acquisition process, the test text is replaced with the property rights text, and the property rights text is replaced with the test text. The distance between two feature vectors refers to the Euclidean distance.
[0093] Step S430: The ratio of the number of difference serial numbers between the test text and the property text to the target number of serial numbers is used as the hash difference value between the test text and the property text; and the infringement index of the test text on the property text is obtained based on the hash difference value and the semantic difference value.
[0094] In a specific implementation of the embodiment of the present invention, the infringement index is expressed by the formula:
[0095]
[0096] Where Qin is the infringement index of the tested text and the property text; is the semantic difference value between the text to be tested and the property text; L is the number of difference serial numbers between the text to be tested and the property text; is the number of target sequence numbers; is the hash difference between the test text and the property text; exp is an exponential function with a natural constant as the base. and The larger the value, the greater the difference in hash value between the test text and the property text and the greater the difference in semantic content, indicating that the text similarity between the test text and the property text is smaller, the possibility that the test text infringes on the property text is smaller, and the infringement index is smaller.
[0097] In a specific implementation of the embodiment of the present invention, the number of target serial numbers is equal to the minimum value of the total number of characters in the hash values of the text to be tested and the property rights text.
[0098] The blockchain may contain multiple property rights. Each property right document described before the current location represents a single property right document on the blockchain. The infringement indicators of the test document against all property rights documents on the blockchain are obtained in the same way as the infringement indicators of the test document against the property right document.
[0099] Based on the infringement index, the test text is judged to see if it infringes the property rights text on the blockchain, and the test text that does not infringe is stored on the blockchain, thereby ensuring the integrity of the property rights registration text data stored on the blockchain. The specific operations are:
[0100] Determine whether the infringement indicators of the text to be tested on all property texts are less than the preset infringement threshold. If so, the text to be tested will not infringe the property text on the blockchain, and the text to be tested will be stored on the node of the blockchain; otherwise, the text to be tested cannot be stored on the node of the blockchain.
[0101] In a specific implementation of the embodiment of the present invention, the preset infringement threshold is set to 0.5.
[0102] So far, the present invention is completed.
[0103] Example 2:
[0104] The present invention also proposes a computer device schematic diagram of a data integrity protection device based on hash function and blockchain technology, please refer to Figure 4 The computer device includes a memory 501, a processor 502, and a computer program 503 stored in the memory 501 and running on the processor 502, wherein when the processor 502 executes the computer program 503, the computer device can execute any one of the data integrity protection methods based on hash function and blockchain technology introduced above.
[0105] In addition, an embodiment of the present application also protects a device, which may include a memory and a processor, wherein the memory stores executable program code, and the processor is used to call and execute the executable program code to perform a data integrity protection method based on hash function and blockchain technology provided by an embodiment of the present application.
[0106] In this embodiment, the device can be divided into functional modules based on the above-described method examples. For example, each functional module can be mapped to a specific functional module, or two or more functions can be integrated into a single processing module. The integrated module can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and represents only a logical functional division. In actual implementation, other division methods may be used.
[0107] In the case of dividing the modules into modules corresponding to their functions, the device may further include a communication module, a signal analysis module, a complexity analysis module, a positioning module, etc. It should be noted that all relevant contents of the various steps involved in the above method embodiment can be referred to the functional description of the corresponding functional modules and will not be repeated here.
[0108] It should be understood that the device provided in this embodiment is used to execute the above-mentioned data integrity protection method based on hash function and blockchain technology, and therefore can achieve the same effect as the above-mentioned implementation method.
[0109] In the case of an integrated unit, the device may include a processing module and a storage module. When the device is applied to a device, the processing module may be used to control and manage the operation of the device. The storage module may be used to support the device in executing mutual program codes, etc.
[0110] The processing module may be a processor or controller that implements or executes the various exemplary logic blocks, modules, and circuits disclosed herein. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor (DSP) and a microprocessor, and the like. The storage module may be a memory.
[0111] Example 3:
[0112] This embodiment also provides a computer-readable storage medium, which stores computer program code. When the computer program code runs on a computer, the computer executes the above-mentioned related method steps to implement a data integrity protection method based on hash function and blockchain technology provided by the above embodiment.
[0113] Example 4:
[0114] This embodiment also provides a computer program product. When the computer program product is run on a computer, it enables the computer to execute the above-mentioned related steps to implement a data integrity protection method based on hash function and blockchain technology provided by the above embodiment.
[0115] Among them, the device, computer-readable storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0116] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0117] It should be noted that the order in which the embodiments of the present invention are described above is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0118] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
Claims
1. A data integrity protection method based on hash function and blockchain technology, characterized in that: The method includes: Obtain the text to be tested and the property rights text stored on the blockchain, and record the text to be tested and the property rights text as analysis text; divide the analysis text into different sentences, and segment the sentences to obtain segmented words; Obtain characteristic words in the analyzed text according to the number of occurrences of each participle in each sentence in the analyzed text and the other participles in the other sentences; Determine the paragraphs in the property text in which each characteristic word in the test text recurs; obtain the overall subject relevance of each characteristic word in the test text relative to the property text based on the number of characters and the number of words in each paragraph in which the characteristic word appears, and the similarity between the entity relationships between the clause in which the characteristic word in the test text appears in the recurring paragraph and the remaining clauses in the property text; Obtaining an infringement index of the test text against the property rights text based on the overall subject correlation difference and textual differences of the characteristic vocabulary of the test text and the property rights text; protecting the data integrity of the blockchain based on the infringement index; The overall subject relevance difference is: based on the characteristic vocabulary of the text to be tested, a hash value of the text to be tested is obtained, and the number of difference numbers in the hash values of the text to be tested and the property text is used as the overall subject relevance difference of the characteristic vocabulary of the text to be tested and the property text; The text difference is as follows: the feature words corresponding to the characters of each difference sequence number of the hash value of the analysis text are converted into word vectors to obtain the feature vectors of the hash value of the analysis text at each difference sequence number, and the average of the distances between the feature vectors of the hash values of the test text and the property text at the same difference sequence number is used as the text difference between the test text and the property text; The hash values of the text to be tested and the property text are different in the characters of each difference sequence number.
2. The data integrity protection method based on hash function and blockchain technology according to claim 1 is characterized in that: The acquiring and analyzing characteristic words in the text includes: Arrange the segmented words of each sentence in the analyzed text in order of position to obtain the segmented word sequence of the corresponding sentence; For each sentence segmentation sequence of the analyzed text, a segmentation in the segmentation sequence is randomly selected and recorded as the target segmentation, and the target segmentation and each subsequent segmentation form a segmentation pair of the target segmentation; in the analyzed text, the proportion of the number of sentences containing each segmentation pair of the target segmentation to the number of sentences containing the target segmentation is used as the characteristic degree of each segmentation pair of the target segmentation; the segmentation pair corresponding to the maximum value of the characteristic degree of the target segmentation is selected and recorded as the initial characteristic vocabulary of the target segmentation; The initial characteristic words of all the segmented words in the analyzed text are recorded as the characteristic words in the analyzed text.
3. The data integrity protection method based on hash function and blockchain technology according to claim 1 is characterized in that: The step of obtaining the relevance of each characteristic word in the text to be tested to the overall subject of the property rights text includes: Obtaining the lexical importance index of each characteristic word in the analysis text based on the similarity of the entity relationship between the clause where each characteristic word is located and the remaining clauses in the analysis text, as well as the distribution position of the remaining clauses in the analysis text; Based on the number of characters and the number of words in the repeated paragraphs of the property text in which the characteristic words in the text to be tested appear, and the lexical importance index of the characteristic words in the text to be tested that appear in the repeated paragraphs, the overall main body relevance of each characteristic word in the text to be tested relative to the property text is obtained.
4. The data integrity protection method based on hash function and blockchain technology according to claim 3 is characterized in that: The acquisition of the important vocabulary index of each characteristic vocabulary in the analysis text includes: Entity relations are extracted for each sentence in the analyzed text to obtain the SPO triples of the corresponding sentences. The ratio of the number of characters in the SPO triple to the total number of characters in the corresponding sentence is recorded as the sentence ratio of the SPO triple. A characteristic word in the analysis text is randomly recorded as an example word, and similar triples of the example word are selected from the SPO triples of each sentence in the analysis text except the sentence where the example word is located. The similar triples have the same first element and the same third element as any SPO triple of the sentence where the example word is located; the average of the sentence ratios of all similar triples of the example word is used as the entity similarity index of the example word; Number the sentences in the analyzed text and calculate the mean of the sequence numbers of the sentences where all similar triples of the example words are located as the position distribution indicator of the example words; The lexical importance index of the example word segmentation is obtained based on the number of similar triples of the example vocabulary, the entity similarity index and the position distribution index; the number of similar triples of the example vocabulary and the entity similarity index are both positively correlated with the lexical importance index, and the position distribution index is negatively correlated with the lexical importance index.
5. The data integrity protection method based on hash function and blockchain technology according to claim 3 is characterized in that: The step of obtaining the overall subject relevance of each characteristic word in the test text to the property text includes: A characteristic word in the test text is randomly selected as the target word, and the ratio of the total number of characters in which the characteristic word in the test text appears in each recurring paragraph of the target word in the property text to the total number of characters in the recurring paragraph is used as the character repetition rate of the target word in each recurring paragraph of the property text; Selecting characteristic words that appear in each recurring paragraph of the target word in the property text in the test text as the focus words of the recurring paragraph; taking the ratio of the mean value of the lexical importance index of the focus words in each recurring paragraph of the target word in the property text to the maximum value of the lexical importance index of the characteristic words in the property text as the text importance of the target word in each recurring paragraph of the property text; Obtaining the local subject relevance of the target word in each recurring paragraph of the property text based on the character repetition degree, the text importance, and the number of times the characteristic words in the test text appear in each recurring paragraph of the target word in the property text; The average of the local subject relevance of the target vocabulary in all recurring paragraphs of the property text is used as the overall subject relevance of the target vocabulary relative to the property text.
6. The data integrity protection method based on hash function and blockchain technology according to claim 1 is characterized in that: The step of obtaining the infringement index of the tested text on the property rights text includes: The ratio of the number of difference serial numbers between the test text and the property text to the target number of serial numbers is used as the hash difference value between the test text and the property text; and the infringement index of the test text on the property text is obtained based on the hash difference value and the text difference between the test text and the property text; The number of target serial numbers is equal to the minimum value of the total number of characters in the hash values of the text to be tested and the property rights text.
7. The data integrity protection method based on hash function and blockchain technology according to claim 1 is characterized in that: The step of obtaining a hash value of the text to be tested based on the characteristic vocabulary of the text to be tested includes: The ratio of the frequency of occurrence of each characteristic word in the test text to the average frequency of occurrence of the characteristic word in all property texts on the blockchain is used as the block association index of each characteristic word in the test text; the block association index of each characteristic word in the test text is negatively correlated, and the product of the mapping result and the overall subject association degree is used as the recognition support index of each characteristic word in the test text; Arranging the recognition support indicators of the characteristic words of the test text in order of their positions to obtain a recognition sequence of the test text; performing a one-dimensional DCT transform on the recognition sequence to obtain a DCT coefficient of each element in the recognition sequence; Calculate the mean of the DCT coefficients of all elements in the recognition sequence as the classification coefficient; set the encoding of the feature vocabulary corresponding to the elements in the recognition sequence whose DCT coefficient is greater than the classification coefficient to 1, and set the encoding of the feature vocabulary corresponding to the elements whose DCT coefficient is less than or equal to the classification coefficient to 0; arrange the encoding of the feature vocabulary of the text to be tested in the order of the position of the feature vocabulary to form a string as the hash value of the text to be tested.
8. The data integrity protection method based on hash function and blockchain technology according to claim 1 is characterized in that: The protection of blockchain data integrity based on the infringement indicators includes: Determine whether the infringement indicators of the text to be tested for all property rights texts are all less than the preset infringement threshold. If so, store the text to be tested on the node of the blockchain. Otherwise, the text to be tested cannot be stored on the node of the blockchain.
9. The data integrity protection method based on hash function and blockchain technology according to claim 2 is characterized in that: Each characteristic word in the text to be tested contained in the recurring paragraph corresponds to two segmented words in the segmented word pair.
Citation Information
Patent Citations
Electronic case duplicate checking method and device based on word segmentation text and computer equipment
CN111814447A
Data processing methods, apparatuses, and devices
US20210326357A1