File information processing method and device based on block chain and electronic equipment

By using text extraction models and word decomposition tree processing machinery to build syntax trees for character matching on the blockchain, the problems of low file processing efficiency and high machine learning model training costs in existing technologies are solved, and efficient and secure file information processing is achieved.

CN120633640APending Publication Date: 2025-09-12INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510771314.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies in file processing have problems such as low file classification and indexing efficiency, difficulty in detecting malicious files, and high cost of training machine learning models, resulting in high overall file information processing costs.

Method used

The target string is extracted by a text extraction model deployed on the blockchain, and a syntax tree is constructed using an independent word decomposition tree processor for character matching, reducing the need to train machine learning models separately for each functional module, and using encryption mechanisms to ensure data security.

Benefits of technology

It reduces the cost of file information processing, improves processing efficiency, and achieves efficient file information matching and secure transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633640A_ABST
    Figure CN120633640A_ABST
Patent Text Reader

Abstract

The invention discloses a file information processing method and device based on a block chain and electronic equipment, and relates to the field of financial science and technology and the technical field of block chains. The method comprises the following steps: extracting a target character string from a target file through a text extraction model deployed on a block chain; after keyword information of the target function module and a target character string are encrypted through a block chain, transmitting the keyword information and the target character string to a word decomposition tree processor; converting the keyword information into a target syntax tree through a constructor; and after the target character string and the target syntax tree are successfully matched, sending the target character string and the keyword information to a target function module for processing. According to the method and the device, the technical problem that the overall file information processing cost is relatively high due to the condition that the model training cost is relatively high because each functional module needs to independently train a machine learning model to complete matching of key information and specific file contents of the functional module in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of financial technology, and specifically, to a file information processing method, device and electronic device based on blockchain. Background Art

[0002] With the widespread application of digital safes in file storage and management, the requirements for their file processing functions are becoming increasingly complex. However, the existing technologies have many shortcomings in achieving these functions.

[0003] File classification and indexing often rely on manual operations or simple pre-processing methods, which are difficult to meet the efficiency and accuracy requirements of large-scale file management. Malicious file detection mainly reads file contents through MIME (Multipurpose Internet Mail Extensions), but this method is inadequate when faced with complex file formats and new attack methods. In addition, derivative functions after file upload, such as batch comparison and sensitive word detection, are usually implemented by independent modules. Most of these modules are based on NLP (Natural Language Processing) and machine learning models. However, training these models requires a large amount of labeled data, which is costly to deploy, and the processing efficiency of a single module is difficult to meet real-time requirements.

[0004] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0005] The embodiments of the present application provide a blockchain-based file information processing method, device, and electronic device to at least solve the technical problem in the prior art that each functional module needs to be individually trained with a machine learning model to match the key information of the functional module with the content of a specific file, thereby being limited by the high cost of model training and resulting in high overall file information processing costs.

[0006] According to one aspect of an embodiment of the present application, a blockchain-based file information processing method is provided, comprising: extracting a target string from a target file through a text extraction model deployed on the blockchain; encrypting keyword information and the target string of a target functional module through the blockchain, and transmitting the encrypted information to a word decomposition tree processor, wherein the word decomposition tree processor is independent of the blockchain, and the word decomposition tree processor comprises: a constructor for constructing a syntax tree and a matcher for matching different character information; converting the keyword information into a target syntax tree through the constructor, wherein each node of the target syntax tree is used to store a character of the keyword information; performing character matching on the target string and the target syntax tree through the matcher; and after the target string and the target syntax tree are successfully matched, sending the target string and the keyword information to the target functional module for processing.

[0007] Optionally, the keyword information is converted into a target syntax tree through a constructor, including: splitting the keyword information according to characters through the constructor to obtain a splitting result; converting the target syntax tree according to the splitting result and the functional module identifier corresponding to each character, wherein the word order relationship between any two characters in the keyword information is represented by the connection relationship between the two nodes corresponding to the two characters in the target syntax tree; and the functional module identifier corresponding to each character is stored in the node corresponding to the character.

[0008] Optionally, the blockchain-based file information processing method further includes: taking the tree branch from the root node of the target syntax tree to the i-th node as the target tree branch; detecting the head and tail length arrays of the character string corresponding to the target tree branch, wherein the head and tail length arrays are used to represent the length of the target substring with the same first character and last character in the character string corresponding to the target tree branch.

[0009] Optionally, after detecting the head and tail length arrays of the string corresponding to the target tree branch, detect whether the value of the head and tail length arrays is greater than 0 or equal to 0; when it is detected that the value of the head and tail length arrays is equal to 0, add a pointer pointing to itself for the i-th node; when it is detected that the value of the head and tail length arrays is greater than 0, add a target pointer to the i-th node according to the value of the head and tail length arrays, wherein the target pointer is used to point to the node j layers back from the i-th node, wherein j is equal to the value of the head and tail length arrays.

[0010] Optionally, detecting the head and tail length arrays of the string corresponding to the target tree branch includes: for the string corresponding to the target tree branch, sorting the characters to disassemble L substrings, where L is an integer greater than 1, and the t+1th substring is obtained by combining the tth substring and a character after the tth substring, and t is a positive integer less than L; when it is detected that the first character and the last character of the tth substring are the same character, recording the length value of the tth substring as the value of the head and tail length array at the tth position, where the head and tail length array is an L-bit array; when it is detected that the first character and the last character of the tth substring are not the same character, setting the value of the head and tail length array at the tth position to 0.

[0011] Optionally, in the process of performing character matching on the target string and the target syntax tree through the matcher, if it is detected that the xth character in the target string is the same as the character on the yth node in the target syntax tree, it is determined that the two characters are matched successfully, and character matching is performed on the x+1th character in the target string and the character on the y+1th node in the target syntax tree, where x and y are both positive integers; if it is detected that the xth character in the target string is not the same as the character on the yth node in the target syntax tree, it is determined that the two characters fail to match, and the character on the node pointed to by the pointer of the yth node is queried, and character matching is performed on the queried character with the xth character in the target string.

[0012] Optionally, in the process of performing character matching on the target string and the target syntax tree through the matcher, if it is detected that the x-th character in the target string and the character on the y-th node in the target syntax tree are not the same, and the character on the node pointed to by the pointer of the y-th node is the first character in the keyword information, then the matching of the x+1-th character in the target string and the first character in the keyword information is performed.

[0013] Optionally, in the process of performing character matching on the target string and the target syntax tree through the matcher, if it is detected that the current character for character matching with the target syntax tree is the last character in the target string, the flag value used to represent the matching position is determined based on the position of the current character in the target string and the character length of the keyword information.

[0014] Optionally, after the target string and the target syntax tree are successfully matched, the target string and keyword information are sent to the target functional module for processing, including: after the target string and the target syntax tree are successfully matched, the target string, keyword information and file information corresponding to the target string are decrypted and sent to the target functional module for processing.

[0015] According to another aspect of an embodiment of the present application, a blockchain-based file information processing device is also provided, including: an extraction unit for extracting a target string from a target file through a text extraction model deployed on the blockchain; a first processing unit for encrypting the keyword information and target string of the target functional module through the blockchain, and transmitting them to a word decomposition tree processor, wherein the word decomposition tree processor is independent of the blockchain, and the word decomposition tree processor includes: a constructor for constructing a syntax tree and a matcher for matching different character information; a conversion unit for converting the keyword information into a target syntax tree through the constructor, wherein each node of the target syntax tree is used to store a character of the keyword information; a second processing unit for performing character matching on the target string and the target syntax tree through the matcher; and a third processing unit for sending the target string and keyword information to the target functional module for processing after the target string and the target syntax tree are successfully matched.

[0016] According to another aspect of an embodiment of the present application, a computer-readable storage medium is further provided, in which a computer program is stored. When the computer program is executed, the device where the computer-readable storage medium is located executes the above-mentioned blockchain-based file information processing method.

[0017] According to another aspect of an embodiment of the present application, an electronic device is also provided, comprising one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by one or more processors, the one or more processors execute the above-mentioned blockchain-based file information processing method.

[0018] According to another aspect of an embodiment of the present application, a computer program product is also provided, wherein the computer program product includes a computer program or instructions, and the computer program or instructions implement the above-mentioned blockchain-based file information processing method when executed by a processor.

[0019] As can be seen from the above content, the present application extracts a target string from a target file through a text extraction model deployed on a blockchain; after encrypting the keyword information and target string of the target functional module through the blockchain, the keyword information and target string are transmitted to a word decomposition tree processor, wherein the word decomposition tree processor is independent of the blockchain, and the word decomposition tree processor includes: a constructor for constructing a syntax tree and a matcher for matching different character information; the keyword information is converted into a target syntax tree through the constructor, wherein each node of the target syntax tree is used to store a character of the keyword information; the target string and the target syntax tree are matched by the matcher; after the target string and the target syntax tree are successfully matched, the target string and keyword information are sent to the target functional module for processing.

[0020] In an embodiment of the present application, there is no need to train a machine learning model separately for each functional module. By using a word decomposition tree processor, the keyword information is converted into a target syntax tree. With the help of the tree structure, the purpose of efficiently matching the keyword information and the target character string using the word decomposition tree processor is achieved, thereby achieving the technical effect of reducing the file information processing cost and improving the processing efficiency. This solves the problem in the prior art that a machine learning model needs to be trained separately for each functional module to complete the matching of the key information of the functional module and the specific file content, thereby being limited by the high cost of model training, resulting in a high overall file information processing cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0022] Figure 1 This is a flowchart of an optional blockchain-based file information processing method according to an embodiment of the present application;

[0023] Figure 2 is a flowchart of another optional blockchain-based file information processing method according to an embodiment of the present application;

[0024] Figure 3 is a schematic diagram of an optional tree based on word decomposition according to an embodiment of the present application;

[0025] Figure 4 This is a schematic diagram of an optional blockchain-based file information processing device according to an embodiment of the present application. DETAILED DESCRIPTION

[0026] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0028] It should also be noted that the information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) collected by this application are information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of relevant data comply with the relevant laws, regulations and standards of the relevant regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse. For example, an interface is set up between this system and relevant users or institutions. Before obtaining relevant information, it is necessary to send an acquisition request to the aforementioned user or institution through the interface, and obtain relevant information after receiving the consent information fed back by the aforementioned user or institution.

[0029] According to an embodiment of the present application, an embodiment of a blockchain-based file information processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0030] Optionally, according to an embodiment of the present application, a blockchain-based file information processing system (hereinafter referred to as the system) is provided as the execution subject of the blockchain-based file information processing method of the embodiment of the present application, wherein the system can be a software system or an embedded system combining software and hardware. Of course, the method execution subject in the embodiment of the present application can also be other forms of execution subjects, such as devices, equipment, etc. Those skilled in the art should know that this application does not specifically limit the specific form of expression of the method execution subject.

[0031] Figure 1 This is a flowchart of an optional blockchain-based file information processing method according to an embodiment of the present application. Figure 1As shown, the method includes the following steps:

[0032] Step S101: extract the target string from the target file using the text extraction model deployed on the blockchain.

[0033] Alternatively, blockchain is a distributed ledger technology with characteristics such as immutability, traceability, and decentralization. In this application, blockchain is used in a digital safe to provide a secure and reliable environment for file storage and management.

[0034] Alternatively, the text extraction model may refer to a model built using NLP technology and OCR (Optical Character Recognition) technology. NLP technology can extract keywords and key sentences from text, while OCR technology can recognize text content in a document.

[0035] Optionally, the target files refer to files uploaded by the user to the digital safe, and these files may include various formats, such as text files, image files, etc.

[0036] Optionally, the target character strings refer to keywords or key sentences extracted from the target file, and these character strings will be used for subsequent matching and processing.

[0037] Optionally, before the target file enters the target functional module (i.e., derivative module) for processing, the system can pre-process the target string by using a text extraction model to filter out keyword information in advance, reduce the amount of data for subsequent processing, and improve overall processing efficiency.

[0038] Step S102: After encrypting the keyword information of the target functional module and the target character string through the blockchain, transmit them to the word decomposition tree processor.

[0039] In step S102, a word decomposition tree processor is independent of the blockchain. The word decomposition tree processor includes: a constructor for constructing a syntax tree and a matcher for matching different character information.

[0040] Optionally, the keyword information of the target functional module is the keywords required by a specific functional module related to target file preprocessing. For example, a sensitive word detection module may require a set of sensitive words as keywords, and a malicious file detection module may require a set of malicious code keywords.

[0041] Optionally, during the preprocessing of the target file, the target character strings extracted through OCR and NLP techniques, as well as the keyword information of the target functional module, are encrypted by an encryption module based on the blockchain's encryption mechanism, thereby protecting the security of this sensitive information during transmission and processing. The system then feeds the encrypted data into the word decomposition tree processor for further processing. The blockchain's encryption mechanism ensures that the data remains secure, even in environments outside the blockchain, and is protected from unauthorized access or tampering.

[0042] Alternatively, by separating the decomposition tree processor from the blockchain, the system can leverage specialized computing resources and optimized algorithms to rapidly process encrypted target strings and keyword information. This reduces the storage and computational burden on the blockchain, allowing it to focus on core functions (such as data recording and verification), thereby improving overall system efficiency. Furthermore, although the decomposition tree processor is independent of the blockchain, encryption mechanisms ensure data security during transmission and processing. The encrypted target strings and keyword information cannot be leaked or tampered with, even in environments outside the blockchain.

[0043] In summary, while blockchain technology offers many advantages, it does have certain limitations when handling tasks like file preprocessing (e.g., processing efficiency, scalability, and storage requirements). Therefore, by separating the decomposition tree processor from the blockchain and encrypting the target string and keyword information, we can improve the overall efficiency and scalability of the system while ensuring the security of the target string and keyword information. This design approach helps overcome the system's over-reliance on blockchain technology and ensures the system's efficiency and security in practical applications.

[0044] Step S103: convert the keyword information into a target syntax tree through a constructor.

[0045] In step S103, each node of the target syntax tree is used to store one character of keyword information.

[0046] Optionally, a constructor typically refers to a function or method used to initialize an object. A constructor can convert keyword information into a target syntax tree. The target syntax tree is a tree-like data structure used to store and organize keyword information. Each node stores a character of the keyword, and the path from the root node to the leaf node represents a complete keyword information.

[0047] Optionally, the target syntax tree stores keyword information in a tree structure, so that each character of each keyword information is stored in a node of the tree. By traversing the tree layer by layer starting from the root node, it is possible to quickly determine whether the target string matches a particular keyword information. During the matching process, if a character does not match, the path can be directly skipped, avoiding the need to compare the entire keyword information one by one. In addition, due to the shared nature of the tree structure, multiple keyword information can share a common prefix, thereby saving storage space.

[0048] Optionally, keyword information can be dynamically added or removed from the target syntax tree without rebuilding the entire tree structure. This allows the system to flexibly adapt to new keyword information needs. Furthermore, this tree structure can accommodate keyword information of varying length and complexity, making it highly versatile.

[0049] Optionally, constructing the target syntax tree may include the following steps: the constructor first creates a root node as the starting point of the target syntax tree. Subsequently, for each keyword information, the constructor starts from the root node and inserts each character into the target syntax tree. If a character already exists on the current path, the next character is inserted along the path; if not, a new node is created. After all characters of a keyword information are inserted into the target syntax tree, the node of the last character is marked as the end of the keyword information insertion.

[0050] For example, assuming the keyword information set is: {"abc", "abd", "abcd"}, the target syntax tree is constructed as follows:

[0051] Step 1. Initialize the root node.

[0052] Step 2. Insert keyword information "abc", including:

[0053] Starting from the root node, create nodes "a", "b", and "c" in sequence, and mark the end of keyword information insertion on the "c" node.

[0054] Step 3. Insert keyword information "abd", including:

[0055] Starting from the root node, continue inserting "d" along the existing path "a" and "b", and mark the end of keyword information insertion on the "d" node.

[0056] Step 4. Insert keyword information "abcd", including:

[0057] Starting from the root node, continue inserting "d" along the existing path "a", "b", and "c", and mark the end of keyword information insertion on the "d" node.

[0058] Step S104: Perform character matching on the target character string and the target syntax tree through a matcher.

[0059] Optionally, a matcher is a component specifically used to compare and match strings by traversing the target syntax tree to check whether characters in the target string match characters in keyword information in the target syntax tree.

[0060] Optionally, the matching process may include the following steps:

[0061] First, the matcher starts from the root node of the target syntax tree and the first character of the target string. Subsequently, the first character of the target string is compared with the character of the root node. Among them, the matcher can adopt the DFS (Depth-First Search) algorithm. Specifically, if the current character of the target string matches the character of the current node of the target syntax tree, it moves to the next character of the target string and continues matching along the corresponding child nodes of the target syntax tree. If the current character of the target string does not match the character of the current node of the target syntax tree, a fallback or jump operation is performed according to the structure of the target syntax tree. For example, the head and tail length array or the failure pointer can be used to quickly jump to the next possible matching starting point. If the matcher reaches a leaf node of the target syntax tree and the leaf node is marked as the end of keyword information insertion, it means that the keyword information is contained in the target string. When all characters of the target string have been checked, the matching process ends.

[0062] Optionally, the structure of the target syntax tree enables the matcher to quickly retrieve keyword information. Through character-by-character comparison and recursive matching, the matcher can complete the matching operation in a shorter time. In addition, by utilizing the shared prefix feature of the target syntax tree, the matcher can avoid multiple matches on repeated prefixes, thereby improving efficiency. Because the target syntax tree can be dynamically updated, the matcher can adapt to new or deleted keyword information without having to rebuild the entire target syntax tree structure. In addition, the matcher can process multiple keyword information at the same time, by traversing the target string once to check whether it contains multiple keyword information.

[0063] Step S105: After the target character string and the target syntax tree are successfully matched, the target character string and keyword information are sent to the target functional module for processing.

[0064] Optionally, after a successful match, the matcher records the location of the matched keyword information within the target string. This can be achieved by recording the index of the target string. The matcher then extracts the matched keyword information and sends it along with the target string to the target functional module. Before sending the match results, the matcher encrypts the target string and keyword information, using an asymmetric encryption algorithm to ensure the security of the target string and keyword information.

[0065] Optionally, Figure 2 is a flowchart of another optional blockchain-based file information processing method according to an embodiment of the present application, such as Figure 2 As shown in the figure, the financial institution's business management platform receives keyword information and information files such as products and documents. This information is sent to a digital safe (also known as a data security guard) for secure processing. In the digital safe, the information is first encrypted by the encryption module, and then the target string is extracted by the TD-IDF extraction model (ie, text extraction model). The encrypted information and the extracted target string can be sent to BAAS (Blockchain as aService, blockchain service platform) through a smart contract. Subsequently, the constructor and matcher in the word decomposition tree processor are used to further process this information. After processing is completed, the decryption module is used to decrypt the encrypted information so that the system can call a large model for more in-depth analysis, such as OCR and NLP. Finally, the derivative function module (ie, the target function module) can provide derivative functions (such as Figure 2 Derivative functions 1, 2, ... n) in the system, such as security checks or data processing, to enhance the overall functionality of the system.

[0066] From the above content, it can be seen that the present application does not require separate training of a machine learning model for each functional module. By using a word decomposition tree processor, the keyword information is converted into a target syntax tree. With the help of the tree structure, the purpose of efficiently matching the keyword information and the target character string using the word decomposition tree processor is achieved, thereby achieving the technical effect of reducing the file information processing cost and improving the processing efficiency. This solves the problem in the prior art that a machine learning model needs to be trained separately for each functional module to complete the matching of the key information of the functional module and the specific file content, which is limited by the high cost of model training and leads to the overall high cost of file information processing.

[0067] In an optional embodiment, keyword information is converted into a target syntax tree through a constructor, including: splitting the keyword information according to characters through the constructor to obtain splitting results; converting to obtain a target syntax tree based on the splitting results and the functional module identifier corresponding to each character, wherein the word order relationship between any two characters in the keyword information is represented by the connection relationship between two nodes corresponding to the two characters in the target syntax tree; and the functional module identifier corresponding to each character is stored in the node corresponding to the character.

[0068] Optionally, the system can first split the keyword information according to characters through a constructor to obtain the splitting results. For example, for the keyword "abc", the splitting result is a character sequence ['a', 'b', 'c']. Then, the system can convert the target syntax tree according to the splitting result and the functional module identifier corresponding to each character. During the construction process, each character is stored as a node of the target syntax tree, and the path from the root node to a leaf node represents a complete string of keyword information. The word order relationship between any two characters in the keyword information is represented by the connection relationship between the two nodes corresponding to the two characters in the target syntax tree. For example, in the keyword information string "abc", the word order relationship between the characters 'a' and 'b' is represented by the parent-child connection relationship between the node 'a' and the node 'b', that is, 'a' is the parent node of 'b'. Finally, the system stores the functional module identifier corresponding to each character on the node corresponding to the character. For example, if the function module identifier corresponding to the keyword information string "abc" is 'F1', the function module identifier 'F1' is stored in the node 'c' (a leaf node of the keyword information string "abc") in the target syntax tree.

[0069] In an optional embodiment, the blockchain-based file information processing method also includes: taking the tree branch from the root node of the target syntax tree to the i-th node as the target tree branch; detecting the head and tail length arrays of the character string corresponding to the target tree branch, wherein the head and tail length arrays are used to represent the length of the target substring with the same first character and last character in the character string corresponding to the target tree branch.

[0070] Optionally, the target tree branch refers to a path from the root node to a specific node, and the character sequence on this path constitutes a string. The head and tail length array is an array used to record the length of the target substring with the same first and last characters in the string corresponding to the target tree branch. Specifically, when constructing a word decomposition-based tree (i.e., a target syntax tree) of the keyword information string, the length j of the substring from the root node to the current node of the tree (assuming the path length is i) with the same first and last characters is calculated, and the head and tail length array [i] = j is assigned.

[0071] For example, Figure 3 is a schematic diagram of an optional tree based on word decomposition according to an embodiment of the present application, such as Figure 3 As shown, the root node (root) is the starting point of the Trie tree (prefix tree) and does not store any characters. Each node represents a character, and the child node represents the next possible character of the character. The end node (leaf node) of the target syntax tree represents a complete keyword information. In some implementations, the leaf node contains additional information, such as counts or tags. In the figure, there is a derived function flag (i.e., function module identifier) ​​under the C node, which indicates a specific operation or function related to the "ABABC" keyword. For example, it means that when the C node is matched, the instructions with function module identifiers "2", "4", and "5" need to be executed. In addition, Figure 3 The following keyword information is given:

[0072] cat:root->c->a->t;

[0073] her:root->h->e->r;

[0074] him:root->h->i->m;

[0075] no:root->n->o;

[0076] nova: root->n->o->v->a;

[0077] root->A->M->E->R->……;

[0078] root->A->M->A->T->……;

[0079] ABAB: root->A->B->A->B;

[0080] ABABC: root->A->B->A->B->C.

[0081] Alternatively, by calculating the length of the first and last characters that are identical, unnecessary comparisons can be skipped during subsequent matching, improving matching efficiency. For example, if a string has the same first and last characters and its length is known, this information can be used to quickly locate the possible starting point of the match during matching.

[0082] In an optional embodiment, after detecting the head and tail length arrays of the string corresponding to the target tree branch, detect whether the value of the head and tail length arrays is greater than 0 or equal to 0; when it is detected that the value of the head and tail length arrays is equal to 0, add a pointer pointing to itself for the i-th node; when it is detected that the value of the head and tail length arrays is greater than 0, add a target pointer for the i-th node according to the value of the head and tail length arrays, wherein the target pointer is used to point to a node j layers back from the i-th node, wherein j is equal to the value of the head and tail length arrays.

[0083] Optionally, after detecting the start and end length arrays of the string corresponding to the target tree branch, the system will determine the value j of the start and end length array [i]. If j is equal to 0, indicating that there is no substring with the same first and last characters on the path from the current node to the root node, then the pointer of the current node pointing to itself is increased. This pointer is used to quickly jump to the current node during the matching process, indicating that no matching fallback path is found; if j is greater than 0, indicating that there is a substring with the same first and last characters, then the pointer of the current node is increased. This pointer points to the node at the -j layer (i.e., fallback j layers) of the current node. This pointer is used to quickly fall back to the possible matching starting point during the matching process.

[0084] Optionally, the system can significantly improve the efficiency of matching target strings with keyword information by using arrays of first and last lengths and corresponding pointers. During the matching process, if a mismatching character is encountered, the system can directly jump to the possible matching starting point using the pointer, rather than backing up layer by layer. Furthermore, in traditional Trie tree matching, if a mismatching character is encountered, the system must backtrack layer by layer until a matching branch is found. However, using arrays of first and last lengths and pointers can reduce this backtracking and directly jump to the possible matching starting point.

[0085] For example, consider a Trie tree containing the keyword string "ababc." For node 'b' (the second 'b'), the values ​​of the first and last length arrays are 2 (the path "ab" from the root node to this node has the same first and last characters, and its length is 2). Therefore, a target pointer can be added to node 'b', pointing to node 'a' two levels below it (the 'a' below the root node).

[0086] In an optional embodiment, the head and tail length arrays of the character string corresponding to the target tree branch are detected, including: for the character string corresponding to the target tree branch, sorting out L substrings according to characters, wherein L is an integer greater than 1, and the t+1th substring is obtained by combining the tth substring and a character after the tth substring, and t is a positive integer less than L; when it is detected that the first character and the last character of the tth substring are the same character, recording the length value of the tth substring as the value of the head and tail length array at the tth position, wherein the head and tail length array is an L-bit array; when it is detected that the first character and the last character of the tth substring are not the same character, setting the value of the head and tail length array at the tth position to 0.

[0087] Optionally, first, the system selects a specific branch (i.e., the target tree branch) from the target syntax tree and obtains the string corresponding to the branch. The obtained string is sorted by character and disassembled into L substrings, where L is an integer greater than 1. This process is performed recursively, that is, each substring is composed of the previous substring plus a new character. Then, for each substring, the system checks whether its first character and last character are the same. If the first character and the last character of the t-th substring are the same, then the length of this substring is recorded as the value of the head and tail length array at the t-th position. The head and tail length array is an array of length L, which is used to store these length values. If the first character and the last character of the t-th substring are not the same, then the value of the head and tail length array at the t-th position is set to 0.

[0088] For example, if Figure 3 As shown on the left side of the figure, assuming the keyword information string is "ABABC", initialize the keyword information string's head and tail length array ComLEN = '[0,0,0,0,0]', and disassemble the keyword information string from beginning to end into layer-by-layer substrings ("A"-"AB"-"ABA"-"ABAB"-"ABABC"), calculate the length of the first and last characters of each layer-by-layer substring, and assign the length to the head and tail length array. The final keyword information string's head and tail length array ComLEN is: [0,0,1,2,0]. Among them, the double pointer algorithm can be used to obtain the length of the first and last characters of the layer-by-layer substring.

[0089] In an optional embodiment, in the process of performing character matching on the target string and the target syntax tree through the matcher, if it is detected that the xth character in the target string is the same as the character on the yth node in the target syntax tree, it is determined that the two characters are matched successfully, and character matching is performed on the x+1th character in the target string and the character on the y+1th node in the target syntax tree, where x and y are both positive integers; if it is detected that the xth character in the target string is not the same as the character on the yth node in the target syntax tree, it is determined that the two characters fail to match, and the character on the node pointed to by the pointer of the yth node is queried, and the queried character is matched with the xth character in the target string.

[0090] Optionally, the matcher of the word decomposition processor will match the target string with the word decomposition-based tree (target syntax tree) one by one. If the current target character matches the keyword character at the current target syntax tree level, the next target character will be located and compared with the keyword information character at the next level. If there is no match, the pointer of the currently matched keyword information character will be used to find the keyword information character at the node at level -j pointed by the pointer, and the current target character will be compared with the keyword information character again.

[0091] For example, assuming the target string is S and the target syntax tree is T, matching starts from the root node and is first initialized: x=1 (the first character of the target string); y=1 (the root node of the target syntax tree).

[0092] Next, compare S[x] and T[y]. If S[x] == T[y], the match is successful, x = x+1, y = y+1, and the system continues to compare the next character. If S[x] != T[y], the match fails, and the system checks whether T[y] has a pointer to another branch node. If there is another branch node, T[y] is updated to the branch node and S[x] is compared with the new T[y] again. If there is no other branch node, the match fails and the result is returned.

[0093] Specifically, if all characters of the target string match the keyword information of the target syntax tree (i.e., x is equal to the length of the target string), the match succeeds. If no further matching can be done at a certain character (i.e., no branch node is available), the match fails.

[0094] In an optional embodiment, during the process of performing character matching on the target string and the target syntax tree through the matcher, if it is detected that the x-th character in the target string and the character on the y-th node in the target syntax tree are not the same, and the character on the node pointed to by the pointer of the y-th node is the first character in the keyword information, then the matching of the x+1-th character in the target string and the first character in the keyword information is performed.

[0095] Optionally, if the -j level falls back to the first keyword information character, a match is performed between the next target character and the first keyword information character.

[0096] For example, suppose the target string is "ade," the path of the target syntax tree is "abd," and the keyword information is "de." The first character of the target string is 'a', and the first node of the target syntax tree is also 'a', resulting in a successful match. The match continues with the second character of the target string, 'd', and the second node of the syntax tree, 'b', but the match fails. After the match fails, the system checks the node pointed to by the pointer of the current node (i.e., 'b') in the target syntax tree.

[0097] Assuming that the 'b' node of the syntax tree has a pointer pointing to 'd' (which matches the first character 'd' of the keyword information), the system will match the next character of the target string ('e') with the first character of the keyword information ('d'). The 'e' of the target string does not match the 'd' of the keyword information, and the match fails. However, the second character of the keyword information is 'e', ​​which happens to match the 'e' of the target string successfully. Therefore, the system can continue to match the subsequent characters of the target string with the keyword information. Finally, the target string "ade" successfully matches the keyword information "de", and the entire matching process is completed.

[0098] In an optional embodiment, during the process of performing character matching on the target string and the target syntax tree through the matcher, if it is detected that the current character for character matching with the target syntax tree is the last character in the target string, then the flag value used to represent the matching position is determined based on the position of the current character in the target string and the character length of the keyword information.

[0099] Optionally, if the current keyword information character is the last one, it means that the system completes a match and records the matching position, which is: the position of the current target character in the target string - the length of the keyword information character + 1.

[0100] For example, if Figure 3 As shown on the left side of the figure, the keyword string is "ABABC", the target string is "ABABABC", the first 4 characters are matched successfully, and the 5th character does not match, then press Figure 3 The pointer to the keyword information character "C" in the left branch is positioned to the third character "A" on the -2 layer, matching the current character (i.e., the fifth character) and the repositioned keyword information character "A". The matching continues in sequence until the keyword information character reaches the last character "C". The matching position is 7-5+1=3, that is, "ABABABC" matches the keyword information character starting from the third character "A".

[0101] Optionally, for some derivative functions (functions corresponding to the target functional module), the layer-by-layer substrings of the corresponding keyword information character strings may have equal head and tail substrings, such as the keyword "ABABC" for malicious file upload verification or sensitive word detection, so constructing an array of the head and tail lengths of the keyword information helps to speed up the search efficiency. In addition, the encrypted keyword information is a random character sequence (a mixture of letters (uppercase or lowercase), numbers or special symbols), the length will be significantly increased, and a large number of character sequence repetitions will be present. The above steps are designed to improve the efficiency of such a large number of repeated encrypted keywords and character string matching.

[0102] In an optional embodiment, after the target string and the target syntax tree are successfully matched, the target string and keyword information are sent to the target functional module for processing, including: after the target string and the target syntax tree are successfully matched, the target string, keyword information and file information corresponding to the target string are decrypted and sent to the target functional module for processing.

[0103] Optionally, each time the target string successfully matches the keyword information, the system can decrypt the keyword information string, the target string and the corresponding file information through the decryption module according to the functional module identifier on the leaf node of the corresponding target syntax tree, and send them to the target functional module corresponding to the functional module identifier for processing.

[0104] For example, the malicious file function module is identified as "2", the encrypted keyword information string is "ABABC", the target string extracted by the text extraction model is "ABABABC", and the file information corresponding to the target string is " / etc / icbc01 / workspace / contact / 00057 / hwxxx.pdf". After decrypting this information, it is sent to the target function module corresponding to the function module identification "2" for further processing.

[0105] Optionally, the function of the word decomposition tree processor can be analogous to a Bloom filter. Since the length of the target string is usually much longer than the keyword information string, this may lead to a decrease in accuracy when matching the target functional module (for example, in the case of malicious file content upload, although some codes match certain keyword information, there is no security risk as a whole). Therefore, similar to the Bloom filter, this method will produce false positive results, and these results need to be sent to a dedicated target functional module for further processing. The filtering effect of the word decomposition tree processor helps to reduce the need for a large number of files to call various target functional modules, thereby avoiding excessive occupation of system resources and improving processing efficiency.

[0106] It should be noted that the working principle of the Bloom filter is based on probability, so there are two possibilities for its results: If the Bloom filter indicates that an element is not in the set, then this conclusion is certain, that is, the element is definitely not in the set. If the Bloom filter shows that an element is in the set, it means that the element may indeed exist in the set, but it may not. In this case, there are two possibilities: one is a true positive (True Positive), that is, the element actually exists in the set; the other is a false positive (False Positive), that is, the element is not actually in the set, but due to the probabilistic nature of the Bloom filter, it mistakenly reports the existence of the element. The design of the Bloom filter ensures that there will be no false negatives, that is, an element that actually exists will not be mistakenly reported as not existing.

[0107] According to another aspect of the embodiment of the present application, a file information processing device based on blockchain is also provided, wherein: Figure 4 is a schematic diagram of an optional blockchain-based file information processing device according to an embodiment of the present application, such as Figure 4 As shown, the blockchain-based file information processing device includes: an extraction unit 401, a first processing unit 402, a conversion unit 403, a second processing unit 404, and a third processing unit 405.

[0108] Optionally, the extraction unit 401 is used to extract the target string from the target file through the text extraction model deployed on the blockchain; the first processing unit 402 is used to encrypt the keyword information and target string of the target functional module through the blockchain, and transmit them to the word decomposition tree processor, wherein the word decomposition tree processor is independent of the blockchain, and the word decomposition tree processor includes: a constructor for constructing a syntax tree and a matcher for matching different character information; the conversion unit 403 is used to convert the keyword information into a target syntax tree through the constructor, wherein each node of the target syntax tree is used to store a character of the keyword information; the second processing unit 404 is used to perform character matching on the target string and the target syntax tree through the matcher; the third processing unit 405 is used to send the target string and keyword information to the target functional module for processing after the target string and the target syntax tree are successfully matched.

[0109] Optionally, the conversion unit 403 includes: a first processing subunit and a conversion subunit. The first processing subunit is configured to split the keyword information by character using a constructor to obtain a split result; the conversion subunit is configured to convert the split result and the functional module identifier corresponding to each character into a target syntax tree, wherein the word order relationship between any two characters in the keyword information is represented by the connection relationship between two nodes corresponding to the two characters in the target syntax tree; and the functional module identifier corresponding to each character is stored in the node corresponding to the character.

[0110] Optionally, the blockchain-based file information processing device further includes: a fourth processing unit and a first detection unit. The fourth processing unit is configured to use a tree branch from the root node of the target syntax tree to the i-th node as a target tree branch; and the first detection unit is configured to detect a head and tail length array of a string corresponding to the target tree branch, wherein the head and tail length array is used to represent the length of a target substring in which the first and last characters of the string corresponding to the target tree branch are identical.

[0111] Optionally, the blockchain-based file information processing device further includes: a second detection unit, a fifth processing unit, and a sixth processing unit. The second detection unit is configured to detect whether the value of the head and tail length arrays is greater than or equal to 0; the fifth processing unit is configured to, upon detecting that the value of the head and tail length arrays is equal to 0, add a pointer to the i-th node pointing to itself; and the sixth processing unit is configured to, upon detecting that the value of the head and tail length arrays is greater than 0, add a target pointer to the i-th node based on the value of the head and tail length arrays, wherein the target pointer is configured to point to a node j layers back from the i-th node, where j is equal to the value of the head and tail length arrays.

[0112] Optionally, the first detection unit includes: a second processing subunit, a third processing subunit, and a fourth processing subunit. The second processing subunit is used to disassemble the string corresponding to the target tree branch into L substrings according to character sorting, wherein L is an integer greater than 1, and the t+1th substring is obtained by combining the tth substring and a character after the tth substring, and t is a positive integer less than L; the third processing subunit is used to record the length value of the tth substring as the value of the first and last length array at the tth position when it is detected that the first character and the last character of the tth substring are the same character, wherein the first and last length array is an L-bit array; the fourth processing subunit is used to set the value of the first and last length array at the tth position to 0 when it is detected that the first character and the last character of the tth substring are not the same character.

[0113] Optionally, the second processing unit 404 includes: a first determining subunit and a second determining subunit. The first determining subunit is used to determine that the xth character in the target string and the character on the yth node in the target syntax tree are the same if the two characters are detected to be the same, and perform character matching on the x+1th character in the target string and the character on the y+1th node in the target syntax tree, where x and y are both positive integers; and the second determining subunit is used to determine that the xth character in the target string and the character on the yth node in the target syntax tree are not the same if the two characters are detected to be different, and query the character on the node pointed to by the pointer of the yth node, and perform character matching on the queried character with the xth character in the target string.

[0114] Optionally, the second processing unit 404 further includes a third determining subunit. The third determining subunit is configured to, if it is detected that the current character being matched with the target syntax tree is the last character in the target string, determine a flag value representing the matching position based on the position of the current character in the target string and the character length of the keyword information.

[0115] Optionally, the third processing unit 405 includes a fifth processing subunit, wherein the fifth processing subunit is configured to decrypt the target string, keyword information, and file information corresponding to the target string and send them to the target functional module for processing after the target string and the target syntax tree are successfully matched.

[0116] According to another aspect of an embodiment of the present application, a computer-readable storage medium is further provided, in which a computer program is stored. When the computer program is executed, the device where the computer-readable storage medium is located executes the above-mentioned blockchain-based file information processing method.

[0117] According to another aspect of an embodiment of the present application, an electronic device is also provided, comprising one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by one or more processors, the one or more processors execute the above-mentioned blockchain-based file information processing method.

[0118] The above-mentioned embodiments or examples disclosed in this application are not exhaustive, but are only illustrations of some embodiments or examples, and are not intended to be specific limitations on the scope of protection disclosed in this application. In the absence of contradiction, each step in a certain embodiment or example in this application can be implemented as an independent example, and the steps can be arbitrarily combined. For example, the solution after removing some steps in a certain embodiment or example can also be implemented as an independent example, and the order of the steps in a certain embodiment or example can be arbitrarily exchanged. In addition, the optional methods or optional examples in a certain embodiment or example can be arbitrarily combined; in addition, the various embodiments or examples can be arbitrarily combined. For example, some or all of the steps in different embodiments or examples can be arbitrarily combined, and a certain embodiment or example can be arbitrarily combined with the optional methods or optional examples of other embodiments or examples.

[0119] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0120] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0121] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0122] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected to achieve the purpose of the present embodiment according to actual needs.

[0123] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0124] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program code.

[0125] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A file information processing method based on blockchain, characterized in that: include: Extract the target string from the target file through the text extraction model deployed on the blockchain; After encrypting the keyword information of the target functional module and the target character string through the blockchain, the encryption is transmitted to a word decomposition tree processor, wherein the word decomposition tree processor is independent of the blockchain and includes: a constructor for constructing a syntax tree and a matcher for matching different character information; Converting the keyword information into a target syntax tree by the constructor, wherein each node of the target syntax tree is used to store one character of the keyword information; Performing character matching on the target string and the target syntax tree by the matcher; After the target character string and the target syntax tree are successfully matched, the target character string and the keyword information are sent to the target function module for processing.

2. The method according to claim 1, characterized in that Converting the keyword information into a target syntax tree by the constructor includes: Splitting the keyword information by characters using the constructor to obtain a split result; The target syntax tree is converted based on the splitting result and the functional module identifier corresponding to each character, wherein the word order relationship between any two characters in the keyword information is represented by the connection relationship between two nodes corresponding to the two characters in the target syntax tree; and the functional module identifier corresponding to each character is stored in the node corresponding to the character.

3. The method according to claim 1, characterized in that The method further comprises: The tree branch from the root node of the target syntax tree to the i-th node is used as the target tree branch; Detect the first and last length arrays of the character string corresponding to the target tree branch, wherein the first and last length arrays are used to represent the length of the target substring with the same first and last characters in the character string corresponding to the target tree branch.

4. The method according to claim 3, characterized in that After detecting the start and end length arrays of the character strings corresponding to the target tree branches, the method further includes: Check whether the value of the first and last length arrays is greater than or equal to 0; When it is detected that the value of the first and last length arrays is equal to 0, a pointer pointing to itself is added to the i-th node; When it is detected that the value of the head and tail length array is greater than 0, a target pointer is added for the i-th node according to the value of the head and tail length array, wherein the target pointer is used to point to the node j layers back from the i-th node, wherein j is equal to the value of the head and tail length array.

5. The method according to claim 3, characterized in that Detecting the first and last length arrays of the strings corresponding to the target tree branches, including: For the string corresponding to the target tree branch, sort the characters and decompose it into L substrings, where L is an integer greater than 1, the t+1th substring is obtained by combining the tth substring and a character following the tth substring, and t is a positive integer less than L; When it is detected that the first character and the last character of the t-th substring are the same character, the length value of the t-th substring is recorded as the value of the first and last length array at the t-th position, wherein the first and last length array is an L-bit array; When it is detected that the first character and the last character of the t-th substring are not the same character, the value at the t-th position of the head and tail length array is set to 0.

6. The method according to claim 3, characterized in that In the process of performing character matching on the target character string and the target syntax tree by the matcher, the method further includes: If it is detected that the xth character in the target string and the character at the yth node in the target syntax tree are the same, then it is determined that the two characters are matched successfully, and character matching is performed on the x+1th character in the target string and the character at the y+1th node in the target syntax tree, where x and y are both positive integers; If it is detected that the x-th character in the target string and the character on the y-th node in the target syntax tree are not the same, it is determined that the two characters fail to match, and the character on the node pointed to by the pointer of the y-th node is queried, and a character match is performed on the queried character with the x-th character in the target string.

7. The method according to claim 6, characterized in that In the process of performing character matching on the target character string and the target syntax tree by the matcher, the method further includes: If it is detected that the xth character in the target string and the character on the yth node in the target syntax tree are not the same, and the character on the node pointed to by the pointer of the yth node is the first character in the keyword information, then a match is performed between the x+1th character in the target string and the first character in the keyword information.

8. The method according to claim 6, characterized in that In the process of performing character matching on the target character string and the target syntax tree by the matcher, the method further includes: If it is detected that the current character matching the target syntax tree is the last character in the target string, a flag value for representing the matching position is determined according to the position of the current character in the target string and the character length of the keyword information.

9. The method according to claim 1, characterized in that After the target character string and the target syntax tree are successfully matched, the target character string and the keyword information are sent to the target function module for processing, including: After the target character string and the target syntax tree are successfully matched, the target character string, the keyword information and the file information corresponding to the target character string are decrypted and sent to the target function module for processing.

10. A file information processing device based on blockchain, characterized in that: include: An extraction unit, configured to extract a target string from a target file using a text extraction model deployed on the blockchain; a first processing unit, configured to encrypt keyword information of a target functional module and the target character string through the blockchain and transmit the encrypted information to a word decomposition tree processor, wherein the word decomposition tree processor is independent of the blockchain and includes: a constructor for constructing a syntax tree and a matcher for matching different character information; a conversion unit, configured to convert the keyword information into a target syntax tree through the constructor, wherein each node of the target syntax tree is used to store one character of the keyword information; A second processing unit, configured to perform character matching on the target string and the target syntax tree through the matcher; The third processing unit is configured to send the target character string and the keyword information to the target function module for processing after the target character string and the target syntax tree are successfully matched.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located executes the blockchain-based file information processing method according to any one of claims 1 to 9.

12. An electronic device, characterized in that: The system comprises one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the blockchain-based file information processing method according to any one of claims 1 to 9.

13. A computer program product, characterized in that The method comprises a computer program or an instruction, which, when executed by a processor, implements the blockchain-based file information processing method of any one of claims 1 to 9.