Method and device for obtaining document cognition and computing equipment

CN121909458APending Publication Date: 2026-04-21BEIJING DATAFIELD TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING DATAFIELD TECHNOLOGY CO LTD
Filing Date
2023-11-13
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

The prior art is difficult to manage documents accurately and efficiently, especially document classification and grading of unstructured data. Existing methods such as keywords, regular expressions and context-based classification methods are not very accurate and rely on a lot of training or Manual participation, inefficient.

Method used

By obtaining the basic events and event combinations, cognitive elements are determined, including cognitive relationship information and cognitive attribute information, updating advanced continuation relationships, and determining and classifying document-level categories based on predefined decisions.

Benefits of technology

It realizes accurate and efficient management of documents, improves the accuracy and efficiency of document classification and grading, reduces manual participation, and is suitable for a variety of document types and scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121909458A_ABST
    Figure CN121909458A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for obtaining document cognition and computing equipment. The method comprises the steps that basic events and / or a combination of the basic events are obtained; determining a cognitive element, the cognitive element comprising cognitive relationship information and cognitive attribute information, the cognitive relationship information being determined according to the basic events and / or the combination of the basic events, and the cognitive attribute information being determined according to the combination of the basic events; the cognitive relationship information comprises one of relationship combination members formed by inter-entity relationships of various document individual attributes and intra-entity relationships of various document individual attributes, and the cognitive attribute information comprises document individual attributes and non-document individual attributes corresponding to the cognitive relationship information; based on the cognitive element, updating the advanced continuation relationship corresponding to the cognitive element and / or the relationship between the advanced continuation relationships; and when the updated advanced continuation relationship and / or the relationship between the advanced continuation relationships accord with a predefined decision, updating the level category of the document corresponding to the updated advanced continuation relationship. According to the method, cognition of the document can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device and computing device for obtaining document recognition

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on June 29, 2023, with application number 202310778334.0 and application name “Method, device and computing device for obtaining document recognition”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of document management, and more particularly, to a method, apparatus, and computing device for obtaining document knowledge. Background Art

[0003] With the continuous improvement of the level of enterprise informatization, data has become one of the important production factors. Enterprises involve a large amount of business data in their business management activities such as industry and services, marketing support, business operations, risk management, information disclosure, and analytical decision-making. These data may contain the company's business secrets, work secrets, and employees' privacy information.

[0004] The key to document management lies in identifying the business and category of a document—that is, document classification and grading. The difficulty in document classification and grading lies in detecting and identifying unstructured data. Data carriers can include, but are not limited to, documents, images, and videos. Therefore, how to manage documents, images, and videos based on their business type and value has become a pressing technical challenge in the industry.

[0005] Related technical solutions for document management include: grading and classifying documents, images, and videos through content analysis; grading and classifying documents, images, and videos through machine learning; grading and classifying documents, images, and videos through context-based classification; and manual grading and classifying documents, images, and videos. Keyword and regular expression-based recognition of structured data has low accuracy and is difficult to implement. Context-based classification is suitable for scenarios where application file formats and categories are strongly correlated, such as automatically classifying documents generated by CAD applications into the design category. However, it is difficult to accurately classify document formats such as doc and pdf, which are weakly correlated with document categories. Artificial intelligence-based classification methods rely on extensive training and are only suitable for a limited number of scenarios, resulting in low overall recognition rates. Manual classification and grading rely on active human participation, significantly impacting work efficiency. Enterprise IT administrators also find it difficult to force users to participate in manual grading and classification, as users often do not actively tag documents, making implementation difficult.

[0006] Therefore, how to manage documents accurately and efficiently has become a technical problem that needs to be solved urgently.

[0007] Summary of the Invention

[0008] The present application provides a method, apparatus, and computing device for obtaining document cognition, which can obtain cognition of a document, thereby managing the document accurately and efficiently.

[0009] In a first aspect, a method for obtaining document cognition is provided, the method comprising: obtaining basic events and / or combinations of basic events; determining cognitive elements, the cognitive elements comprising cognitive relationship information and cognitive attribute information, wherein the cognitive relationship information is determined based on the basic events and / or combinations of the basic events, and the cognitive relationship information comprises one of the relationship combination members consisting of inter-entity relationships of various document individual attributes and intra-entity relationships of various document individual attributes, and the cognitive attribute information comprises document individual attributes and non-document individual attributes corresponding to the cognitive relationship information; based on the cognitive elements, updating the high-level continuation relationship and / or the relationship between high-level continuation relationships corresponding to the cognitive elements, the high-level continuation relationship being the data updated by the cognitive element combination associated with the corresponding data set and\or the cognitive element combination based on rules; when the updated high-level continuation relationship and / or the relationship between high-level continuation relationships meets a predefined decision, updating the level category of the document corresponding to the updated high-level continuation relationship, wherein the predefined decision comprises high-level continuation relationship features and their corresponding level categories.

[0010] In combination with the first aspect, in certain implementations of the first aspect, based on the cognitive element, the high-level continuation relationship and / or the relationship between high-level continuation relationships corresponding to the cognitive element are updated, including: updating the high-level continuation relationship and / or the relationship between high-level continuation relationships corresponding to the cognitive element according to the type of the cognitive element, wherein the type of the cognitive element is determined according to the cognitive relationship information and / or the creation information of the high-level continuation relationship corresponding to the cognitive element, and the creation information indicates that the high-level continuation relationship corresponding to the cognitive element has been created or the high-level continuation relationship corresponding to the cognitive element has not been created.

[0011] In combination with the first aspect, in certain implementations of the first aspect, the type of the cognitive element conforms to the creation member cognitive element of the advanced continuation relationship, and updating the advanced continuation relationship corresponding to the cognitive element according to the type of the cognitive element includes: creating a new advanced continuation relationship corresponding to the cognitive element.

[0012] In combination with the first aspect, in certain implementations of the first aspect, the type of the cognitive element conforms to the continuation member cognitive element of the advanced continuation relationship, and updating the advanced continuation relationship corresponding to the cognitive element according to the type of the cognitive element includes: continuing the advanced continuation relationship corresponding to the cognitive element.

[0013] In combination with the first aspect, in certain implementations of the first aspect, the high-level continuation relationship includes at least one of the following relationships: document mirror recognition, document entity flow relationship, document content flow relationship, and folder recognition.

[0014] In combination with the first aspect, in certain implementations of the first aspect, based on the cognitive element, the high-level continuation relationship and / or the relationship between the high-level continuation relationships corresponding to the cognitive element are updated, including: based on the folder attributes of the document operated by the cognitive element, updating the folder cognition corresponding to the folder attribute, wherein the folder cognition is a type of the high-level continuation relationship.

[0015] In combination with the first aspect, in certain implementations of the first aspect, the cognitive element is a cognitive element that expresses a document mirror entity derivation relationship or a cognitive element that expresses a document network transmission relationship. Based on the cognitive element, the high-level continuation relationship and / or the relationship between high-level continuation relationships corresponding to the cognitive element are updated, including: associating the document operated by the cognitive element and the copy generated by the cognitive element with the same document entity flow relationship, wherein the document entity flow relationship is a type of the high-level continuation relationship.

[0016] In combination with the first aspect, in certain implementations of the first aspect, the document of the cognitive meta-operation has a unique corresponding document mirror cognition, and the document mirror cognition is a type of the high-level continuation relationship.

[0017] In combination with the first aspect, in certain implementations of the first aspect, determining the cognitive element includes: based on the storage location of the document operated by the cognitive element, selecting at least one of the following methods to store the extended attributes of the document operated by the cognitive element, and the extended attributes include addressing data of high-level continuation relationships: embedding the extended attributes into extensible attributes in the file format; storing the extended attributes for modifying the document content; encrypting and encapsulating the extended attributes and the main text file into one file; storing the extended attributes in the extensible attribute part of the file system; storing the extended attributes in a predefined database or file.

[0018] In combination with the first aspect, in certain implementations of the first aspect, based on the cognitive element, the high-level continuation relationship and / or the relationship between the high-level continuation relationships corresponding to the cognitive element are updated, including: when the cognitive element indicates that the destination document is at risk of losing the addressing data of the high-level continuation relationship of the source document, the high-level continuation relationship addressing data of the destination document is not lost.

[0019] In combination with the first aspect, in certain implementations of the first aspect, in the step of updating the high-level continuation relationship corresponding to the cognitive element and / or the relationship between the high-level continuation relationships, at least an individual identifier stored in the document extension attribute is used, and the individual identifier is used to determine the correspondence between the document and the high-level continuation relationship.

[0020] In combination with the first aspect, in certain implementations of the first aspect, the individual identifier includes at least one of the following: a document identifier, a document mirror identifier, and an identifier that enables a one-to-one mapping of a document to a document mirror.

[0021] In combination with the first aspect, in some implementations of the first aspect, the relationship between the high-level continuation relationships includes the degree of the relationship between the high-level continuation relationships and the direction of the relationship between the high-level continuation relationships, wherein the direction of the relationship between the high-level continuation relationships is determined based on the source and destination of the cognitive relationship information.

[0022] In combination with the first aspect, in certain implementations of the first aspect, based on the cognitive element, the high-level continuation relationship corresponding to the cognitive element and / or the relationship between the high-level continuation relationships are updated, including: modifying the high-level continuation relationship corresponding to the cognitive element and the relationship between the high-level continuation relationships based on the cognitive relationship information and the non-document individual attribute.

[0023] In combination with the first aspect, in certain implementations of the first aspect, the high-level continuation relationship corresponding to the cognitive element and the relationship between the high-level continuation relationships are modified based on the cognitive relationship information and the non-document individual attribute, including: updating the document entity flow relationship or document content flow relationship corresponding to the document operated by the cognitive element based on the cognitive relationship information, and updating the relationship between the document entity flow relationship or document content flow relationship corresponding to the document operated by the cognitive element and the document entity flow relationship or document content flow relationship corresponding to the documents operated by other cognitive elements based on the non-document individual attribute; or updating the folder cognition corresponding to the folder attribute of the document operated by the cognitive element based on the folder attribute, and updating the relationship between the folder cognition corresponding to the folder attribute and the folder cognition corresponding to other folder attributes based on the cognitive relationship information, wherein the non-document individual attribute includes the folder attribute.

[0024] In combination with the first aspect, in certain implementations of the first aspect, the method is applied to hierarchically classify the document.

[0025] In conjunction with the first aspect, in certain implementations of the first aspect, before updating the level category of the document corresponding to the updated high-level continuation relationship, the method further includes: determining an attribute feature of the high-level continuation relationship, the attribute feature including at least one of the following: a boundary feature, an importance of the high-level continuation relationship, a folder attribute corresponding to the high-level continuation relationship, an order of subject attributes corresponding to the high-level continuation relationship, and an application combination method of the high-level continuation relationship;

[0026] When the updated high-level continuation relationship and / or the relationship between the high-level continuation relationships meets the predefined decision, updating the level category of the document corresponding to the updated high-level continuation relationship includes: when the attribute characteristics of the high-level continuation relationship or the combination of the attribute characteristics of the high-level continuation relationship meets the first decision subset in the predefined decision, updating the level category of the document corresponding to the updated high-level continuation relationship, the first decision subset including the attribute characteristics of the high-level continuation relationship and its corresponding level category and\or the combination of the attribute characteristics of the high-level continuation relationship and its corresponding level category.

[0027] In conjunction with the first aspect, in certain implementations of the first aspect, when an attribute feature of the high-level continuation relationship or a combination of attribute features of the high-level continuation relationship satisfies a first decision subset of predefined decisions, updating the level category of the document corresponding to the updated high-level continuation relationship includes:

[0028] If the boundary characteristics and subject attribute sequence of the advanced continuation relationship meet the first decision subset, determining that the advanced continuation relationship meeting the attribute feature combination is a predefined level category, and / or determining that the document corresponding to the advanced continuation relationship meeting the attribute feature combination is a predefined level category, and the subject attribute sequence includes a time series of subject attributes; or

[0029] If the boundary characteristics and importance of the high-level continuation relationship meet the first decision subset, determining that the high-level continuation relationship meeting the attribute feature combination is a predefined level category, and / or determining that the document corresponding to the high-level continuation relationship meeting the attribute feature combination is a predefined level category; or

[0030] If the boundary characteristics and application combination mode of the advanced continuation relationship meet the first decision subset, determining that the advanced continuation relationship meeting the attribute feature combination is a predefined level category, and / or determining that the document corresponding to the advanced continuation relationship meeting the attribute feature combination is a predefined level category; or

[0031] If the combination of the boundary feature, the application combination mode, and the importance of the advanced continuation relationship meets the first decision subset, determining that the advanced continuation relationship meeting the attribute feature combination is a predefined level category, and / or determining that the document corresponding to the advanced continuation relationship meeting the attribute feature combination is a predefined level category; or

[0032] If the combination of the boundary characteristics of the advanced continuation relationship, the application combination method and the folder attributes corresponding to the advanced continuation relationship meets the first decision subset, the advanced continuation relationship that meets the attribute feature combination is determined to be a predefined level category, and\or, the document corresponding to the advanced continuation relationship that meets the attribute feature combination is determined to be a predefined level category.

[0033] In combination with the first aspect, in certain implementations of the first aspect, when an updated high-level continuation relationship and / or a relationship between high-level continuation relationships conforms to a predefined decision, updating the level category of the document corresponding to the updated high-level continuation relationship includes: when a folder attribute of the updated high-level continuation relationship conforms to a second decision subset in the predefined decision, updating the level category of the document corresponding to the updated high-level continuation relationship, the second decision subset including the folder attribute of the high-level continuation relationship and its corresponding level category.

[0034] In combination with the first aspect, in certain implementations of the first aspect, when the updated high-level continuation relationship and / or the relationship between high-level continuation relationships conforms to a predefined decision, the level category of the document corresponding to the updated high-level continuation relationship is updated, including: when the relationship between the first high-level continuation relationship and the second high-level continuation relationship, the attribute characteristics of the first high-level continuation relationship, and the combination of the level category of the second high-level continuation relationship conform to the third decision subset in the predefined decision, the level category of the document corresponding to the first high-level continuation relationship is updated, the first high-level continuation relationship is the high-level continuation relationship corresponding to the document of the cognitive meta-operation, and the third decision subset includes the attribute characteristics of the first high-level continuation relationship, the relationship between the first high-level continuation relationship and the second high-level continuation relationship, the combination of the level categories of the second high-level continuation relationship and its corresponding level category.

[0035] In combination with the first aspect, in certain implementations of the first aspect, the cognitive element is an entity relationship category cognitive element of the individual attributes of the document. Based on the cognitive element, the high-level continuation relationship and / or the relationship between the high-level continuation relationships corresponding to the cognitive element are updated, including: creating a high-level continuation relationship corresponding to the destination document, and at least part of the information in the cognitive attributes of the high-level continuation relationship corresponding to the destination document comes from the cognitive attributes of the high-level continuation relationship corresponding to the source document.

[0036] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: in response to a change in the level category of the first high-level continuation relationship, using at least one of the following methods to determine whether to change the level category of the third high-level continuation relationship or the level category of the document corresponding to the third high-level continuation relationship, wherein the first high-level continuation relationship is the high-level continuation relationship corresponding to the document of the cognitive meta-operation: comparing the credibility between the determination method of the level category of the first high-level continuation relationship and the determination method of the level category of the third high-level continuation relationship; determining whether the credibility of the determination method of the level category of the first high-level continuation relationship is greater than a threshold; determining whether the determination method of the level category of the first high-level continuation relationship is a predefined determination method; comparing the direction of the relationship between the first high-level continuation relationship and the third high-level continuation relationship; comparing the importance of the first high-level continuation relationship and the importance of the third high-level continuation relationship.

[0037] In combination with the first aspect, in certain implementations of the first aspect, based on the cognitive element, the high-level continuation relationship corresponding to the cognitive element and / or the relationship between the high-level continuation relationships is updated, including: in response to the document operated by the cognitive element, the high-level continuation relationship corresponding to the document is determined based on the addressing data of the high-level continuation relationship stored in the extended attributes of the document.

[0038] In combination with the first aspect, in certain implementations of the first aspect, based on the cognitive element, the high-level continuation relationship corresponding to the cognitive element and / or the relationship between the high-level continuation relationships are updated, including: creating the high-level continuation relationship corresponding to the cognitive element and the relationship between the newly created high-level continuation relationship and the existing high-level continuation relationship.

[0039] In combination with the first aspect, in some implementations of the first aspect, the method further includes: transmitting the document and / or the level category of the document to a predefined network device based on the level category of the document and the network data of the document transmitted by the network.

[0040] In combination with the first aspect, in some implementations of the first aspect, the method further includes: based on the relationship between the high-level continuation relationships, aggregating multiple high-level continuation relationships of different categories to generate an audit drawing, where the audit drawing is used to represent the distribution of documents on the user device.

[0041] In combination with the first aspect, in some implementations of the first aspect, the method further includes: generating an audit drawing based on multiple high-level continuation relationships, where the audit drawing is used to represent a description of the working status of the subject attributes.

[0042] Optionally, the subject attribute may be a user, a user group, a device, or a device group.

[0043] In combination with the first aspect, in some implementations of the first aspect, the method further includes: determining whether the user behavior is normal based on at least one high-level continuation relationship and a predefined knowledge determination model or a data-driven model.

[0044] In combination with the first aspect, in certain implementations of the first aspect, the method is applied to controlling access to a document, where the access to a document includes: opening a file, writing to a file, deleting a file, changing file permissions, changing file attributes, opening an email message in a mail store, deleting an email in a mail store, retrieving a document from a document management system, synchronizing with a predefined document of a document management system, storing a document in a document management system, or any act of accessing a document or a document repository.

[0045] In combination with the first aspect, in certain implementations of the first aspect, during the document access process, the level category of the document, the storage location of the application corresponding to the document, and the individual identifier in the document's extended attributes are used.

[0046] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: controlling access to or use of the document based on a combination of the following three: a first high-level continuation relationship corresponding to the document, the relationship between the first high-level continuation relationship and a fourth high-level continuation relationship, and the level category of the fourth high-level continuation relationship.

[0047] In combination with the first aspect, in certain implementations of the first aspect, a method for determining the level category of the fourth high-level continuation relationship is a predefined level category determination method, or the fourth high-level continuation relationship is a combination of predefined cognitive meta-categories.

[0048] In a second aspect, a device for obtaining document cognition is provided, the device comprising: an acquisition module for obtaining a basic event and / or a combination of basic events; a processing module for determining a cognitive element, the cognitive element comprising cognitive relationship information and cognitive attribute information, wherein the cognitive relationship information is determined based on the basic event and / or the combination of basic events, and the cognitive relationship information comprises one of relationship combination members consisting of inter-entity relationships of various document individual attributes and intra-entity relationships of various document individual attributes, and the cognitive attribute information comprises document individual attributes and non-document individual attributes corresponding to the cognitive relationship information; the processing module is further configured to update, based on the cognitive element, a high-level continuation relationship and / or a relationship between high-level continuation relationships corresponding to the cognitive element, the high-level continuation relationship being data updated by a combination of cognitive elements associated with a corresponding data set and / or the cognitive element combination based on a rule;

[0049] The processing module is further configured to update the level category of the document corresponding to the updated high-level continuation relationship when the updated high-level continuation relationship and / or the relationship between the high-level continuation relationships meets a predefined decision, wherein the predefined decision includes high-level continuation relationship features and their corresponding level categories.

[0050] In combination with the second aspect, in certain implementations of the second aspect, the processing module is specifically used to: update the high-level continuation relationship and / or the relationship between high-level continuation relationships corresponding to the cognitive element according to the type of the cognitive element, wherein the type of the cognitive element is determined based on the cognitive relationship information and / or the creation information of the high-level continuation relationship corresponding to the cognitive element, and the creation information indicates that the high-level continuation relationship corresponding to the cognitive element has been created or the high-level continuation relationship corresponding to the cognitive element has not been created.

[0051] In combination with the second aspect, in some implementations of the second aspect, the type of the cognitive element conforms to the creation member cognitive element of the advanced continuation relationship, and the processing module is specifically used to: create a new advanced continuation relationship corresponding to the cognitive element.

[0052] In combination with the second aspect, in some implementations of the second aspect, the type of the cognitive element conforms to the continuation member cognitive element of the high-level continuation relationship, and the processing module is specifically used to: continue the high-level continuation relationship corresponding to the cognitive element.

[0053] In combination with the second aspect, in some implementations of the second aspect, the high-level continuation relationship includes at least one of the following relationships: document mirror recognition, document entity flow relationship, document content flow relationship, and folder recognition.

[0054] In combination with the second aspect, in certain implementations of the second aspect, the processing module is specifically used to: update the folder cognition corresponding to the folder attribute based on the folder attribute of the document of the cognitive meta-operation, wherein the folder cognition is a type of the high-level continuation relationship.

[0055] In combination with the second aspect, in certain implementations of the second aspect, the cognitive element is a cognitive element that expresses a document mirror entity derivation relationship or a cognitive element that expresses a document network transmission relationship, and the processing module is specifically used to: associate the document operated by the cognitive element and the copy generated by the cognitive element with the same document entity flow relationship, wherein the document entity flow relationship is a type of the high-level continuation relationship.

[0056] In combination with the second aspect, in certain implementations of the second aspect, the document of the cognitive meta-operation has a unique corresponding document mirror cognition, and the document mirror cognition is a type of the high-level continuation relationship.

[0057] In combination with the second aspect, in certain implementations of the second aspect, the processing module is specifically used to: based on the storage location of the document of the cognitive meta-operation, select at least one of the following methods to store the extended attributes of the document of the cognitive meta-operation, and the extended attributes include addressing data of high-level continuation relationships: embedding the extended attributes into the extensible attributes in the file format; storing the extended attributes for modifying the document content; encrypting and encapsulating the extended attributes and the main text file into one file; storing the extended attributes in the extensible attributes part of the file system; storing the extended attributes in a predefined database or file.

[0058] In combination with the second aspect, in some implementations of the second aspect, the processing module is specifically configured to: when the cognitive element indicates that the destination document is at risk of losing the addressing data of the high-level continuation relationship of the source document, ensure that the high-level continuation relationship addressing data of the destination document is not lost.

[0059] In combination with the second aspect, in certain implementations of the second aspect, the processing module uses at least an individual identifier stored in the document extension attribute when executing the step of updating the high-level continuation relationship corresponding to the cognitive element and / or the relationship between the high-level continuation relationships. The individual identifier is used to determine the correspondence between the document and the high-level continuation relationship.

[0060] In combination with the second aspect, in certain implementations of the second aspect, the individual identifier includes at least one of the following: a document identifier, a document mirror identifier, and an identifier that enables a one-to-one mapping of a document to a document mirror.

[0061] In combination with the second aspect, in certain implementations of the second aspect, the relationship between the high-level continuation relationships includes the degree of the relationship between the high-level continuation relationships and the direction of the relationship between the high-level continuation relationships, wherein the direction of the relationship between the high-level continuation relationships is determined based on the source and destination of the cognitive relationship information.

[0062] In combination with the second aspect, in some implementations of the second aspect, the processing module is specifically used to: modify the high-level continuation relationship corresponding to the cognitive element and the relationship between the high-level continuation relationships based on the cognitive relationship information and the non-document individual attribute.

[0063] In combination with the second aspect, in certain implementations of the second aspect, the processing module is specifically used to: update the document entity flow relationship or document content flow relationship corresponding to the document of the cognitive meta-operation based on the cognitive relationship information, and update the relationship between the document entity flow relationship or document content flow relationship corresponding to the document of the cognitive meta-operation and the document entity flow relationship or document content flow relationship corresponding to the documents of other cognitive meta-operations based on the non-document individual attribute; or update the folder cognition corresponding to the folder attribute based on the folder attribute of the document of the cognitive meta-operation, and update the relationship between the folder cognition corresponding to the folder attribute and the folder cognition corresponding to other folder attributes based on the cognitive relationship information, wherein the non-document individual attribute includes the folder attribute.

[0064] In combination with the second aspect, in some implementations of the second aspect, the device is used to hierarchically classify the document.

[0065] In combination with the second aspect, in certain implementations of the second aspect, the processing module is further used to: determine the attribute characteristics of the high-level continuation relationship, and the attribute characteristics include at least one of the following: boundary characteristics, the importance of the high-level continuation relationship, the folder attribute corresponding to the high-level continuation relationship, the subject attribute order corresponding to the high-level continuation relationship, and the application combination method of the high-level continuation relationship; wherein, the processing module is specifically used to: when the attribute characteristics of the high-level continuation relationship or the combination of the attribute characteristics of the high-level continuation relationship meets the first decision subset in the predefined decision, update the level category of the document corresponding to the updated high-level continuation relationship, and the first decision subset includes the attribute characteristics of the high-level continuation relationship and its corresponding level category and\or the combination of the attribute characteristics of the high-level continuation relationship and its corresponding level category.

[0066] In conjunction with the second aspect, in some implementations of the second aspect, the processing module is specifically configured to:

[0067] If the boundary characteristics and subject attribute sequence of the advanced continuation relationship meet the first decision subset, determining that the advanced continuation relationship meeting the attribute feature combination is a predefined level category, and / or determining that the document corresponding to the advanced continuation relationship meeting the attribute feature combination is a predefined level category, and the subject attribute sequence includes a time series of subject attributes; or

[0068] If the boundary characteristics and importance of the high-level continuation relationship meet the first decision subset, determining that the high-level continuation relationship meeting the attribute feature combination is a predefined level category, and / or determining that the document corresponding to the high-level continuation relationship meeting the attribute feature combination is a predefined level category; or

[0069] If the boundary characteristics and application combination mode of the advanced continuation relationship meet the first decision subset, determining that the advanced continuation relationship meeting the attribute feature combination is a predefined level category, and / or determining that the document corresponding to the advanced continuation relationship meeting the attribute feature combination is a predefined level category; or

[0070] If the combination of the boundary feature, the application combination mode, and the importance of the advanced continuation relationship meets the first decision subset, determining that the advanced continuation relationship meeting the attribute feature combination is a predefined level category, and / or determining that the document corresponding to the advanced continuation relationship meeting the attribute feature combination is a predefined level category; or

[0071] If the combination of the boundary characteristics of the advanced continuation relationship, the application combination method and the folder attributes corresponding to the advanced continuation relationship meets the first decision subset, the advanced continuation relationship that meets the attribute feature combination is determined to be a predefined level category, and\or, the document corresponding to the advanced continuation relationship that meets the attribute feature combination is determined to be a predefined level category.

[0072] In combination with the second aspect, in certain implementations of the second aspect, the processing module is specifically used to: when the folder attributes of the updated high-level continuation relationship meet the second decision subset in the predefined decision, update the level category of the document corresponding to the updated high-level continuation relationship, and the second decision subset includes the folder attributes of the high-level continuation relationship and its corresponding level category.

[0073] In combination with the second aspect, in certain implementations of the second aspect, the processing module is specifically used to: when the relationship between the first high-level continuation relationship and the second high-level continuation relationship, the attribute characteristics of the first high-level continuation relationship, and the combination of the level categories of the second high-level continuation relationship meet the third decision subset in the predefined decision, update the level category of the document corresponding to the first high-level continuation relationship, the first high-level continuation relationship is the high-level continuation relationship corresponding to the document of the cognitive meta-operation, and the third decision subset includes the attribute characteristics of the first high-level continuation relationship, the relationship between the first high-level continuation relationship and the second high-level continuation relationship, the combination of the level categories of the second high-level continuation relationship and its corresponding level categories.

[0074] In combination with the second aspect, in certain implementations of the second aspect, the cognitive element is an entity relationship category cognitive element of the individual attributes of the document, and the processing module is specifically used to: create a high-level continuation relationship corresponding to the new destination document, and at least part of the information in the cognitive attributes of the high-level continuation relationship corresponding to the destination document comes from the cognitive attributes of the high-level continuation relationship corresponding to the source document.

[0075] In combination with the second aspect, in certain implementations of the second aspect, the processing module is further used to: in response to a change in the level category of the first high-level continuation relationship, use at least one of the following methods to determine whether to change the level category of the third high-level continuation relationship or the level category of the document corresponding to the third high-level continuation relationship, wherein the first high-level continuation relationship is the high-level continuation relationship corresponding to the document of the cognitive meta-operation: compare the credibility between the determination method of the level category of the first high-level continuation relationship and the determination method of the level category of the third high-level continuation relationship; determine whether the credibility of the determination method of the level category of the first high-level continuation relationship is greater than a threshold; determine whether the determination method of the level category of the first high-level continuation relationship is a predefined determination method; compare the direction of the relationship between the first high-level continuation relationship and the third high-level continuation relationship; compare the importance of the first high-level continuation relationship and the importance of the third high-level continuation relationship.

[0076] In combination with the second aspect, in some implementations of the second aspect, the processing module is specifically used to: in response to the document of the cognitive meta-operation, determine the high-level continuation relationship corresponding to the document based on the addressing data of the high-level continuation relationship stored in the extended attributes of the document.

[0077] In conjunction with the second aspect, in certain implementations of the second aspect, the processing module is specifically configured to: create a high-level continuation relationship corresponding to the cognitive element and a relationship between the newly created high-level continuation relationship and the existing high-level continuation relationship.

[0078] In combination with the second aspect, in certain implementations of the second aspect, the processing module is further used to: transmit the document and / or the level category of the document to a predefined network device based on the level category of the document and the network data of the document transmitted by the network.

[0079] In combination with the second aspect, in some implementations of the second aspect, the processing module is further used to: based on the relationship between the high-level continuation relationships, collect multiple high-level continuation relationships of different categories to generate an audit drawing, which is used to represent the distribution of documents on the user device.

[0080] In combination with the second aspect, in some implementations of the second aspect, the processing module is further used to: generate an audit drawing based on multiple high-level continuation relationships, where the audit drawing is used to represent a description of the working status of the subject attributes.

[0081] In combination with the second aspect, in some implementations of the second aspect, the processing module is further used to: determine whether the user behavior is normal based on at least one high-level continuation relationship and a predefined knowledge determination model or a data-driven model.

[0082] In combination with the second aspect, in certain implementations of the second aspect, the device is applied to control access to a document, where access to the document includes: opening a file, writing to a file, deleting a file, changing file permissions, changing file attributes, opening an email message in a mail storage, deleting an email in a mail storage, retrieving a document from a document management system, synchronizing with a predefined document of a document management system, storing a document in a document management system, or any act of accessing a document or a document repository.

[0083] In combination with the second aspect, in certain implementations of the second aspect, the level category of the document, the storage location of the application corresponding to the document, and the individual identifier in the document's extended attributes are used during the access process of the document.

[0084] In combination with the second aspect, in certain implementations of the second aspect, the processing module is further used to: control access to or use of the document based on a combination of the following three: the first high-level continuation relationship corresponding to the document, the relationship between the first high-level continuation relationship and the fourth high-level continuation relationship, and the level category of the fourth high-level continuation relationship.

[0085] In combination with the second aspect, in certain implementations of the second aspect, a method for determining a level category of the fourth high-level continuation relationship is a predefined level category determination method, or the fourth high-level continuation relationship is a combination of predefined cognitive meta-categories.

[0086] In a third aspect, a method for obtaining cognition of a document is provided, the method comprising: obtaining a basic event; determining a cognitive element based on the basic event and / or a combination of the basic events, the cognitive element comprising cognitive relationship information determined by the basic event and cognitive attribute information determined by the basic event; determining at least one high-level continuation relationship related to the cognitive element based on the cognitive relationship information determined by the basic event and the cognitive attribute information determined by the basic event; and selectively updating the at least one high-level continuation relationship and / or the relationship between high-level continuation relationships based on the at least one high-level continuation relationship.

[0087] In combination with the third aspect, in certain implementations of the third aspect, the cognitive element further includes addressing data between the document and the high-level continuation relationship.

[0088] In combination with the third aspect, in certain implementations of the third aspect, the method also includes: determining cognitive attribute information of the basic event based on at least two of the following information: application attributes, device attributes, user attributes, path attributes, document extension attributes, and time attributes of the basic event and / or the combination of the basic events.

[0089] In combination with the third aspect, in certain implementations of the third aspect, the method further includes: determining the cognitive relationship information determined by the basic event based on at least one of the following information: intra-entity relationships of individual document attributes, inter-entity relationships of individual document attributes, and the inter-entity relationships of individual document attributes include document mirror entity derivation relationships and document network transmission relationships.

[0090] In combination with the third aspect, in certain implementations of the third aspect, the method further includes: storing addressing data between the document and the advanced continuation relationship according to at least one of the following information: document extended metadata storing addressing data between the document and the advanced continuation relationship, a predefined database or file storing addressing data between document attributes and the advanced continuation relationship.

[0091] In combination with the third aspect, in certain implementations of the third aspect, the method further includes: when there is a risk of failing to accumulate the cognitive attributes of the advanced continuation relationship corresponding to the source document based on the cognitive relationship information of the basic event, determining a method for storing addressing data of the advanced continuation relationship corresponding to the source document in a document extension attribute based on the location attribute of the destination document.

[0092] In combination with the third aspect, in certain implementations of the third aspect, the method further includes: determining the cognitive element of the intra-entity relationship of the individual attributes of the document based on the basic event and / or the combination of the basic events, and updating the cognitive attribute information of the high-level continuation relationship corresponding to the source document.

[0093] In combination with the third aspect, in certain implementations of the third aspect, the method further includes: determining the cognitive element of the entity relationship of the individual attributes of the document based on the basic event and / or the combination of the basic events, creating the cognition of the new document mirror and\or updating the high-level continuation relationship corresponding to the source document.

[0094] In conjunction with the third aspect, in certain implementations of the third aspect, the cognition of the created new document image is determined based on a decision by combining the cognition of the source document image and the cognitive elements.

[0095] In combination with the third aspect, in certain implementations of the third aspect, the cognitive element of the document mirror entity derivative relationship is determined based on the basic event and / or the combination of the basic events, and based on the decision, the cognitive attribute information of the at least one high-level continuation relationship is selectively updated.

[0096] In combination with the third aspect, in certain implementations of the third aspect, the cognitive element of the document being transmitted over the network is determined based on the basic event and / or a combination of the basic events, and based on the decision, the cognitive attribute information of at least one high-level continuation relationship is selectively updated.

[0097] In combination with the third aspect, in certain implementations of the third aspect, the method further includes: updating a class of high-level continuation relationships based on the cognitive elements of the document mirror entity derived relationship and the cognitive elements of the document network transmitted relationship.

[0098] In combination with the third aspect, in certain implementations of the third aspect, the at least one high-level continuation relationship and the relationship between the high-level continuation relationships are updated simultaneously, and the relationship between the high-level continuation relationships includes a degree of relationship and / or a direction of relationship.

[0099] In combination with the third aspect, in certain implementations of the third aspect, the method further includes: creating new category cognition elements according to the high-level continuation relationship, maintaining the order of the category cognition elements according to the high-level continuation relationship, and updating the high-level continuation relationship determined by the addressing data according to the addressing data of the high-level continuation relationship.

[0100] In combination with the third aspect, in certain implementations of the third aspect, the method further includes: if the cognitive element is any one of the cognitive element category combinations predefined for the advanced continuation relationship, updating the advanced continuation relationship determined by the addressing data according to the addressing data of the advanced continuation relationship.

[0101] In combination with the third aspect, in certain implementations of the third aspect, the method is applied to hierarchically classify the document.

[0102] In combination with the third aspect, in certain implementations of the third aspect, the method further includes: if the order of subject attributes conforms to the decision, determining that the first high-level continuation relationship is a predefined level category and\or determining that the document corresponding to the first high-level continuation relationship is a predefined level category; or if the combination of multiple stored category cognitive elements conforms to the decision, determining that the first high-level continuation relationship is a predefined level category and\or determining that the document corresponding to the first high-level continuation relationship is a predefined level category.

[0103] In combination with the third aspect, in certain implementations of the third aspect, the method further includes: if the level category of the first high-level continuation relationship changes, and the method for determining the level category of the first high-level continuation relationship is a predefined level category determination method, based on the relationship between high-level continuation relationships, updating the level category of the second high-level continuation relationship and\or the level category of the second document corresponding to the second high-level continuation relationship; the method for determining the level category of the first high-level continuation relationship is a predefined level category determination method, which is determined based on at least one of the following information: the order of the subject attributes of the high-level continuation relationship conforms to the decision, the combination of multiple category cognitive elements stored in the high-level continuation relationship conforms to the decision, manual labeling, and level categories obtained from third-party applications.

[0104] In combination with the third aspect, in certain implementations of the third aspect, the method further includes: if the level category of the first high-level continuation relationship changes, and the method for determining the level category of the first high-level continuation relationship is a predefined level category determination method, based on the relationship between high-level continuation relationships, updating the level category of the second high-level continuation relationship and\or the level category of the second document corresponding to the second high-level continuation relationship, wherein at least the first type of cognitive attribute is used in the determination process of the level category of the first high-level continuation relationship, and at least the second type of cognitive attribute is used in the determination process of the relationship between the high-level continuation relationships.

[0105] In conjunction with the third aspect, in certain implementations of the third aspect, in response to a change in the level category of the first high-level continuation relationship, the level category of the second document or the second high-level continuation relationship is changed.

[0106] In combination with the third aspect, in certain implementations of the third aspect, the level category of the second document or the second high-level continuation relationship is determined by comparing the cognitive meta-category combination of the first high-level continuation relationship and the cognitive meta-category combination of the second high-level continuation relationship.

[0107] In combination with the third aspect, in certain implementations of the third aspect, in the step of updating the level category of the second document and / or the second high-level continuation relationship based on the relationship between the high-level continuation relationships, the relationship between the high-level continuation relationships includes the degree of the relationship and / or the direction of the relationship.

[0108] In combination with the third aspect, in certain implementations of the third aspect, a level category of the second document or the second high-level continuation relationship is determined based on the first high-level continuation relationship.

[0109] In combination with the third aspect, in certain implementations of the third aspect, the method further includes: based on the addressing data between the first document and the advanced continuation relationship and the advanced continuation relationship cognitive attributes, in response to the cognitive element of the first document, changing the level category of the second document or the second advanced continuation relationship.

[0110] In combination with the third aspect, in certain implementations of the third aspect, the method further includes: based on the addressing data between the first document and the advanced continuation relationship, the cognitive attributes of the advanced continuation relationship, and the degree of relationship between the advanced continuation relationships, in response to the cognitive element of the first document, changing the level category of the second document or the second advanced continuation relationship.

[0111] In combination with the third aspect, in certain implementations of the third aspect, the method further includes: determining a combination of the level category of the document and network data transmitted by the document over the network, and transmitting the combination to a predefined network device.

[0112] In combination with the third aspect, in certain implementations of the third aspect, the method is applied to generate a document family distribution graph.

[0113] In combination with the third aspect, in certain implementations of the third aspect, the method further includes: based on the relationship between the high-level continuation relationships, aggregating multiple high-level continuation relationships of different categories to generate an audit drawing, where the audit drawing is used to represent the distribution of documents on the user device.

[0114] In combination with the third aspect, in certain implementations of the third aspect, the method is applied to controlling access to a document, where the document access includes opening a file, writing to a file, deleting a file, changing file permissions, changing file attributes, opening an email message in a mail store, deleting an email in a mail store, retrieving a document from a document management system, storing a document in a document management system, or any act of accessing a document or a document repository.

[0115] In combination with the third aspect, in certain implementations of the third aspect, the method further includes: controlling access to or use of the document based on a combination of: a first high-level continuation relationship corresponding to the document, the relationship between the first high-level continuation relationship and the second high-level continuation relationship, and a level category of the second high-level continuation relationship.

[0116] In conjunction with the third aspect, in certain implementations of the third aspect, the method for determining the level category of the second high-level continuation relationship may be a predefined level category determination method or a combination of determining the second high-level continuation relationship as predefined cognitive meta-categories.

[0117] In a fourth aspect, a device for obtaining document cognition is provided, comprising: an acquisition module and a processing module, wherein the acquisition module is used to obtain a basic event; the processing module is used to determine a cognitive element based on the basic event and / or a combination of the basic events, the cognitive element including cognitive relationship information determined by the basic event and cognitive attribute information determined by the basic event; determining at least one high-level continuation relationship related to the cognitive element based on the cognitive relationship information and the cognitive attribute information; and selectively updating the at least one high-level continuation relationship and / or the relationship between high-level continuation relationships based on the at least one high-level continuation relationship.

[0118] In combination with the fourth aspect, in certain implementations of the fourth aspect, the cognitive element further includes addressing data between the document and the high-level continuation relationship.

[0119] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is also used to determine the cognitive attribute information based on at least two of the following information: application attributes, device attributes, user attributes, path attributes, document extension attributes, and time attributes of the basic event and / or the combination of the basic events.

[0120] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is also used to determine the cognitive relationship information based on at least one of the following information: intra-entity relationships of individual document attributes, inter-entity relationships of individual document attributes, and the inter-entity relationships of individual document attributes include document mirror entity derivation relationships and document network transmission relationships.

[0121] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is further used to store the addressing data between the document and the advanced continuation relationship according to at least one of the following information: document extended metadata stores the addressing data between the document and the advanced continuation relationship, and a predefined database or file stores the addressing data between document attributes and the advanced continuation relationship.

[0122] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is also used to determine, based on the location attributes of the destination document, a document extended attribute to store the addressing data method of the advanced continuation relationship corresponding to the source document, the cognitive relationship information determined by the basic event, and the cognitive attribute information determined by the basic event when there is a risk of not being able to accumulate the advanced continuation relationship cognitive attributes corresponding to the source document according to the cognitive relationship information determined by the basic event.

[0123] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is also used to determine the cognitive element of the intra-entity relationship of the individual attributes of the document based on the basic event and / or the combination of the basic events, update the cognition of the source document mirror, or update the cognitive attribute information of the high-level continuation relationship corresponding to the source document.

[0124] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is also used to determine the cognitive elements of the entity relationship of the individual attributes of the document based on the basic event and / or the combination of the basic events, create the cognition of the new document mirror, and\or update the cognitive attribute information of the high-level continuation relationship corresponding to the source document.

[0125] In conjunction with the fourth aspect, in certain implementations of the fourth aspect, the cognition of the created new document image is determined based on a decision by combining the cognition of the source document image and the cognition element.

[0126] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is specifically used to: determine the cognitive element of the document mirror entity derivative relationship based on the basic event and / or the combination of the basic events, and based on the decision, selectively update the cognitive attribute information of at least one high-level continuation relationship.

[0127] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is specifically used to: determine the cognitive element of the document being transmitted over the network based on the basic event and / or the combination of the basic events, and based on the decision, selectively update the cognitive attribute information of at least one high-level continuation relationship.

[0128] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is further used to update a class of high-level continuation relationships based on the cognitive elements of the document mirror entity derivative relationship and the cognitive elements of the document network transmission relationship.

[0129] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is specifically used to: simultaneously update at least one of the high-level continuation relationships and the relationship between the high-level continuation relationships, where the relationship between the high-level continuation relationships includes the degree of the relationship and / or the direction of the relationship.

[0130] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is also used to create new category cognition elements according to the high-level continuation relationship, maintain the order of category cognition elements according to the high-level continuation relationship, and update the high-level continuation relationship determined by the addressing data according to the addressing data of the high-level continuation relationship.

[0131] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is further used to update the high-level continuation relationship determined by the addressing data according to the addressing data of the high-level continuation relationship if the cognitive element is any one of the cognitive element category combinations predefined for the high-level continuation relationship.

[0132] In combination with the fourth aspect, in certain implementations of the fourth aspect, the device is used to hierarchically classify the document.

[0133] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is further used to determine that the first high-level continuation relationship is a predefined level category and\or determine that the document corresponding to the first high-level continuation relationship is a predefined level category if the order of subject attributes conforms to the decision; or the processing module is further used to determine that the first high-level continuation relationship is a predefined level category and\or determine that the document corresponding to the first high-level continuation relationship is a predefined level category if the combination of multiple stored category cognitive elements conforms to the decision.

[0134] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is also used to update the level category of the second high-level continuation relationship and\or the level category of the second document corresponding to the second high-level continuation relationship based on the relationship between high-level continuation relationships if the level category of the first high-level continuation relationship changes, and the method for determining the level category of the first high-level continuation relationship is a predefined level category determination method; the method for determining the level category of the first high-level continuation relationship is a predefined level category determination method, which is determined based on at least one of the following information: the order of the subject attributes of the high-level continuation relationship conforms to the decision, the combination of multiple category cognitive elements stored in the high-level continuation relationship conforms to the decision, manual labeling, and level categories obtained from third-party applications.

[0135] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is further used to update the level category of the second high-level continuation relationship and\or the level category of the second document corresponding to the second high-level continuation relationship based on the relationship between high-level continuation relationships if the level category of the first high-level continuation relationship changes, and the method for determining the level category of the first high-level continuation relationship is a predefined level category determination method, wherein at least the first type of cognitive attribute is used in the determination process of the level category of the first high-level continuation relationship, and at least the second type of cognitive attribute is used in the determination process of the relationship between the high-level continuation relationships.

[0136] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is specifically configured to: change the level category of the second document or the second high-level continuation relationship in response to a change in the level category of the first high-level continuation relationship.

[0137] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is specifically used to: determine the level category of the second document or the second high-level continuation relationship by comparing the cognitive meta-category combination of the first high-level continuation relationship and the cognitive meta-category combination of the second high-level continuation relationship.

[0138] In combination with the fourth aspect, in certain implementations of the fourth aspect, in the step of updating the level category of the second document and / or the second high-level continuation relationship based on the relationship between the high-level continuation relationships, the relationship between the high-level continuation relationships includes the degree of the relationship and / or the direction of the relationship.

[0139] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is specifically configured to: determine a level category of the second document or the second high-level continuation relationship based on the first high-level continuation relationship.

[0140] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is further used to change the level category of the second document or the second high-level continuation relationship in response to the cognitive element of the first document based on the addressing data between the first document and the high-level continuation relationship and the cognitive attributes of the high-level continuation relationship.

[0141] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is further used to change the level category of the second document or the second high-level continuation relationship in response to the cognitive element of the first document based on the addressing data between the first document and the high-level continuation relationship, the cognitive attributes of the high-level continuation relationship, and the degree of relationship between the high-level continuation relationships.

[0142] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is further used to determine a combination of the level category of the document and the network data of the document transmitted by the network, and transmit the combination to a predefined network device.

[0143] In combination with the fourth aspect, in certain implementations of the fourth aspect, the device is used to generate a document family distribution graph.

[0144] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is further used to generate an audit drawing based on the relationship between the high-level continuation relationships, by collecting multiple high-level continuation relationships of different categories, and the audit drawing is used to represent the distribution of documents on the user device.

[0145] In combination with the fourth aspect, in certain implementations of the fourth aspect, the device is applied to control access to documents, where the document access includes opening a file, writing to a file, deleting a file, changing file permissions, changing file attributes, opening an email message in a mail storage, deleting an email in a mail storage, retrieving a document from a document management system, storing a document in a document management system, or any act of accessing a document or a document repository.

[0146] In combination with the fourth aspect, in certain implementations of the fourth aspect, the access or use of the control document is based on a combination of the following three: the first high-level continuation relationship corresponding to the document, the relationship between the first high-level continuation relationship and the second high-level continuation relationship, and the level category of the second high-level continuation relationship.

[0147] In a fifth aspect, a computing device is provided, comprising a processor and a memory, and optionally, an input / output interface. The processor is configured to control the input / output interface to send and receive information, the memory is configured to store a computer program, and the processor is configured to retrieve and execute the computer program from the memory, causing the computing device to perform the method of the first aspect or any possible implementation of the first aspect, or causing the computing device to perform the method of the third aspect or any possible implementation of the third aspect.

[0148] Optionally, the processor may be a general-purpose processor, which may be implemented in hardware or software. When implemented in hardware, the processor may be a logic circuit, an integrated circuit, or the like; when implemented in software, the processor may be a general-purpose processor implemented by reading software code stored in a memory, which may be integrated into the processor or located independently of the processor.

[0149] In a sixth aspect, a chip is provided, which obtains instructions and executes the instructions to implement the method in the above-mentioned first aspect and any one of the implementations of the first aspect, or to implement the method in the above-mentioned third aspect and any one of the implementations of the third aspect.

[0150] Optionally, as an implementation method, the chip includes a processor and a data interface, and the processor reads instructions stored in the memory through the data interface to execute the method in the above-mentioned first aspect and any one of the implementation methods of the first aspect, or executes the above-mentioned third aspect and any one of the implementation methods of the third aspect.

[0151] Optionally, as an implementation method, the chip may also include a memory, in which instructions are stored, and the processor is used to execute the instructions stored on the memory. When the instructions are executed, the processor is used to execute the method in the first aspect and any one of the implementation methods of the first aspect, or execute the method in the third aspect and any one of the implementation methods of the third aspect.

[0152] In the seventh aspect, a computer program product comprising instructions is provided, which, when executed by a computing device, causes the computing device to execute the method as described in the first aspect and any one of the implementations of the first aspect, or to execute the method as described in the third aspect and any one of the implementations of the third aspect.

[0153] In an eighth aspect, a computer-readable storage medium is provided, comprising computer program instructions. When the computer program instructions are executed by a computing device, the computing device executes the method as described in the first aspect and any one of the implementations of the first aspect, or executes the method as described in the third aspect and any one of the implementations of the third aspect.

[0154] By way of example, these computer-readable storage media include, but are not limited to, one or more of the following: read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), Flash memory, electrically EPROM (EEPROM), and a hard drive.

[0155] Optionally, as an implementation manner, the above-mentioned storage medium may specifically be a non-volatile storage medium. BRIEF DESCRIPTION OF THE DRAWINGS

[0156] FIG1 is a schematic flowchart of a method for obtaining document recognition provided in an embodiment of the present application.

[0157] FIG2 is a schematic block diagram of a basic-level event acquirer 200 .

[0158] FIG3 is a schematic diagram of an audit drawing of a document provided in an embodiment of the present application.

[0159] FIG4 is a schematic flow chart of a method for including cognitive elements of an event in which a document is newly created, provided in an embodiment of the present application.

[0160] FIG5 is a schematic flowchart of a method for maintaining a correspondence between a document and a high-level continuation relationship (MultiRelationNode) provided in an embodiment of the present application.

[0161] FIG6 is a schematic flowchart of a method for including cognitive elements of an event in which a document is uploaded, provided in an embodiment of the present application.

[0162] FIG7 is a schematic flowchart of a method for determining a FileFlowNode with an associated relationship provided in an embodiment of the present application.

[0163] FIG8 is a schematic flowchart of a method for mutually influencing document classification results between family trees provided by an embodiment of the present application.

[0164] FIG9 is a schematic block diagram of an apparatus 900 for obtaining document recognition provided in an embodiment of the present application.

[0165] FIG10 is a schematic diagram of the architecture of a computing device 1500 provided in an embodiment of the present application. DETAILED DESCRIPTION

[0166] The technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.

[0167] The technical solution in this application will be described below with reference to the accompanying drawings.

[0168] This application will present various aspects, embodiments, or features around systems including multiple devices, components, modules, etc. It should be understood and appreciated that each system may include additional devices, components, modules, etc., and / or may not include all of the devices, components, modules, etc. discussed in conjunction with the figures. Furthermore, combinations of these aspects may also be used.

[0169] Additionally, in the embodiments of this application, words such as "exemplary" and "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in this application as "exemplary" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of the word "exemplary" is intended to present concepts in a concrete manner.

[0170] In the embodiments of the present application, “corresponding” and “relevant” may sometimes be used interchangeably. It should be noted that when the distinction between them is not emphasized, the meanings they intend to express are consistent.

[0171] The business scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. A person skilled in the art will appreciate that, with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are equally applicable to similar technical problems.

[0172] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0173] In this application, "at least one" means one or more, and "more" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: including the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.

[0174] The embodiments of the present application provide a method for obtaining document cognition, which can manage documents while also grading and classifying documents with high accuracy based on their sensitivity, value, and security requirements.

[0175] It should be noted that the document recognition may include the level category of the high-level continuation relationship corresponding to the document, and the level category group to which the high-level continuation relationship corresponding to the document belongs. The level category may include business attributes.

[0176] It should be noted that the method for obtaining document recognition provided in the embodiment of the present application can also be applied to scenarios where unstructured data such as pictures and videos are graded and classified according to the sensitivity, value and security requirements of the pictures and videos.

[0177] It should be noted that the original attributes of a document refer to the inherent attributes of the unstructured data at the application level, such as the document path, document size, and file name of the document itself, which are stored in the PC file system.

[0178] It should be noted that decision-based determination means that it can be determined based on the evaluation results of a knowledge-driven model or a data-driven model. For the sake of clarity, rules are used in the embodiments of this application to express selective control based on rule evaluation results.

[0179] It should be noted that selective determination refers to determination based on the evaluation results of a knowledge-driven model or a data-driven model. For the sake of clarity, rules are used in the embodiments of this application to express selective control based on the rule evaluation results.

[0180] In the embodiments of this application, "family tree" and "document family", "FileFlowNode.ActionBusiness", "FileFlowNode", "document entity flow relationship", and "document entity flow relationship cognition" can sometimes be used interchangeably. It should be pointed out that when the distinction between them is not emphasized, the meanings they intend to express are consistent.

[0181] In the embodiments of the present application, the entity relationships of individual document attributes and the document mirror entity change relationships can sometimes be used interchangeably. It should be noted that when the distinction between them is not emphasized, the meanings they intend to express are consistent.

[0182] In the embodiments of the present application, the intra-entity relationship of individual document attributes and the document mirror entity maintenance relationship can sometimes be used interchangeably. It should be noted that when the distinction between them is not emphasized, the meanings they intend to express are consistent.

[0183] In the embodiments of the present application, "EntityMirror.ActionBusiness", "document mirror" and "document mirror cognition" can sometimes be used interchangeably. It should be pointed out that when the distinction between them is not emphasized, the meanings they intend to express are consistent.

[0184] Figure 1 is a schematic flow chart of a method for obtaining document recognition provided by an embodiment of the present application. As shown in Figure 1, the method may include steps 110-130, and steps 110-130 are described in detail below.

[0185] Step 110: Determine the cognitive element.

[0186] The cognitive element includes cognitive relationship information and cognitive attribute information. The cognitive relationship information is determined based on basic events and / or combinations of basic events, and includes one of the relationship combination members consisting of inter-entity relationships and intra-entity relationships of various document-specific attributes. The cognitive attribute information includes document-specific attributes and non-document-specific attributes corresponding to the cognitive relationship information.

[0187] Alternatively, since the addressing data of a high-level continuation relationship can be determined jointly by cognitive relationship information and cognitive attribute information, in another possible implementation, a cognitive element can be understood as a combination of cognitive relationship information, cognitive attribute information, and the addressing data of a high-level continuation relationship. To better illustrate the cognitive element, the present embodiment introduces the concept of a document operation event.

[0188] In an embodiment of the present application, the document operation event may have one or more cognitive relationship information, which is obtained by combining a basic-level event or multiple basic-level events on the client. The combination of the multiple basic-level events may include, but is not limited to, merging multiple basic-level events and associating multiple basic-level events.

[0189] The aforementioned cognitive relationship information can be determined by basic events and / or combinations of basic events. The cognitive relationship information includes one member of a relationship combination consisting of inter-entity relationships and intra-entity relationships of various document attributes. In other words, the cognitive relationship information includes one type of relationship in the relationship combination consisting of inter-entity relationships and intra-entity relationships of various document attributes.

[0190] In other words, if a type of relationship determined by basic events and / or combinations of basic events belongs to a type of relationship in a combination of multiple types of relationships between entities that express individual attributes of documents and multiple types of relationships within entities that express individual attributes of documents, cognitive relationship information is generated.

[0191] In the embodiments of the present application, the inter-entity relationship of individual document attributes represents the change relationship of individual document attributes between entities, and in some embodiments, it can be represented by EntityMirrorChanged. The intra-entity relationship of individual document attributes represents the change relationship of individual document attributes within an entity, and in some embodiments, it can be represented by EntityMirrorRemain.

[0192] In some embodiments, the relationship combination may include: intra-entity relationships of document individual attributes, inter-entity relationships of document individual attributes, and the inter-entity relationships of document individual attributes include document mirror entity derivation relationships, document network transmission relationships, etc.

[0193] A document operation event can have one or more different categories of cognitive relationship information.

[0194] As an example, the cognitive relationship information determined by basic events and / or combinations of basic events is expressed from different dimensions.

[0195] It should be understood that in order to vividly express cognitive relationship information and cognitive attribute information, the document operation events exemplified in the embodiments of the present application may include but are not limited to: a file being created, saved as a new file, a file being copied, a file being burned, a file being compressed, a file being decompressed, a file being archived, a file being uploaded, a file being downloaded, a file being uploaded and obtaining network information, a file being downloaded and obtaining network information, a file being moved, a file being renamed, a file being edited, a file being read-only, content being pasted to the clipboard, content being dragged, etc.

[0196] In one example, basic events and / or combinations of basic events determine the entity relationships EntityMirrorChanged between individual attributes of various types of documents, including but not limited to: a file being created, saved as a new file, a file being copied, a file being burned, a file being compressed, a file being decompressed, a file being archived, a file being uploaded, a file being downloaded, a file being uploaded and obtaining network information, and a file being downloaded and obtaining network information.

[0197] In another example, basic events and / or combinations of basic events determine the entity relationship EntityMirrorRemain of various document individual attributes, including but not limited to: file moved, file renamed, file edited, file read-only, content pasted to the clipboard, and content dragged.

[0198] Another example, the basic events and / or the combination of basic events determine that, in accordance with the decision, a document is transmitted over the network (DataInMotion) relationship is obtained. The basic events and / or the combination of basic events that express this relationship may include but are not limited to: a file is uploaded, a file is downloaded, a file is uploaded and network information is obtained, a file is downloaded and network information is obtained, etc.

[0199] In another example, the basic events and / or the combination of basic events determine that, in accordance with the decision, a document mirror entity derived (FileDerived) relationship is obtained. The basic events and / or the combination of basic events that express this relationship may include but are not limited to: the file is copied, the file is burned, the file is compressed, the file is decompressed, the file is archived, and saved as a new file.

[0200] Another example, the basic events and\or the combination of basic events determine that, in accordance with the decision, the document is obtained as being in use (DataInUse) relationship. The basic events and / or the combination of basic events that express this relationship may include but are not limited to: the file is edited, the file is read-only, the content is pasted to the clipboard, the content is dragged, etc.

[0201] Another example, the basic events and\or the combination of basic events determine that, in accordance with the decision, the document is obtained as a static storage state (DataAtRest) relationship. The basic events and / or the combination of basic events that express this relationship may include but are not limited to: a file is newly created, a file is moved, a file is renamed; a file is copied, a file is burned, a file is compressed, a file is decompressed, a file is archived, etc.

[0202] Another example, the basic event and\or the combination of basic events determine that, in accordance with the decision, a content reference (CopyContent) relationship is obtained. The basic event and / or the combination of basic events that express this relationship may include but is not limited to the examples in this application: content is pasted to the clipboard, content is dragged, and inserted into a file.

[0203] A cognito element that expresses information about a certain type of cognitive relationship is called a cognito element of that category, for example, a cognito element of the FileDerive category derived from a document mirror entity.

[0204] The above-mentioned cognitive attribute information includes document-specific attributes and / or non-document-specific attributes corresponding to cognitive relationship information determined by basic events and / or combinations of basic events.

[0205] Document individual attributes include original document individual attributes (such as document absolute path, document hash), and individual identifiers stored in document extended attributes (such as document identifier FileID). FileID is used as the individual identifier stored in the extended attribute, and the new document entity has a new FileID. The document extended metadata container stores FileID, and the individual identifier can be determined by the original document individual attributes, such as using the hash value of the source document as the FileID value. For example, when different users download the same file from the server, for example, 10 copies are obtained. Since the file does not have a corresponding FileID when it is just downloaded, the hash value of the document is used as the FileID value of the document. The advantage of this is that since these 10 copies have the same hash value of the source document of the download event, the destination document of the download event is associated with the same source document FileID value. It can be understood as ten copies of the same document.

[0206] Non-document individual attributes may include: application attributes (AppBusiness), user attributes (UserBusiness), device attributes (DeviceBusiness), path attributes (FolderBusiness), time attributes (TimeBusiness), etc.

[0207] It should be noted that a cognize set node refers to a data set based on a rule of at least one cognize. A high-level continuation relation is a subset of a cognize set node.

[0208] The high-level continuation relationship is the data updated by the combination of cognitive elements of the associated corresponding data set and\or the combination of cognitive elements based on rules. The associated corresponding data set refers to the storage location of the high-level continuation relationship whose addressing attribute corresponds to the cognitive element, and the combination of cognitive elements includes one or more types of cognitive elements. The cognitive relationship information category of the cognitive elements in the combination of cognitive elements belongs to one of the predefined cognitive relationship category combinations of the high-level continuation relationship. The cognitive element combination is the data updated based on the rules, and the updated data includes the cognitive attributes of the updated high-level continuation relationship, and the cognitive attributes include the importance of the high-level continuation relationship, the order of the subject attributes of the high-level continuation relationship, etc. Further, the cognitive attributes include the level category of the high-level continuation relationship. The update includes creating new data based on the decision, and updating existing data based on the decision. In other words, the high-level continuation relationship is the data updated by one or more types of cognitive elements of the associated corresponding data set and\or the one or more types of cognitive elements based on rules. In the embodiment of the present application, the description involving "cognitive element set node" can also be replaced by "high-level continuation relationship".

[0209] There are many ways to combine various types of advanced continuation relationships, which are not specifically limited in the embodiments of this application.

[0210] Advanced continuation relationships can be summarized using MultiRelationNode, and specific advanced continuation relationships can have specific names, such as FileFlowNode.ActionBusines.

[0211] Addressing data for higher-level continuation relationships: The combination of cognitive relationship information and cognitive attribute information can determine the addressing data for the higher-level continuation relationship of the destination document of a document operation event. In one embodiment, the document's extended metadata stores the addressing data for the higher-level continuation relationship, for example, the addressing data between the document and the corresponding MultiRelationNode. In another embodiment, based on the cognitive relationship information and cognitive attribute information, the addressing relationship between the document's original attributes and the higher-level continuation relationship is determined in a predefined location, such as a database.

[0212] The document ID is the information used to identify the EntityMirror identifier stored in the document's extended metadata. New documents have new IDs. The document's extended metadata container stores the document ID, which actually establishes an entity correspondence between the document's original attributes and the document's mirror. An optional method maintains the relationship between the document's original attributes and the document's mirror, addressing the document mirror based on the ID. In response to events where the document entity does not change, the document ID and EntityMirror.ID remain unchanged. An optional method updates the document ID in response to events where the document entity does not change, and determines the EntityMirror.ID based on the updated document ID. In other words, regardless of whether the document ID changes or not, the core is to maintain the original attributes of the document - the addressing of the document mirror.

[0213] A cognate that contains a cognitive relationship defined by a base event is a cognate of that category. For example, a cognate that contains the document mirror entity derivative relationship FileDerived is a cognate of the document mirror entity derivative relationship FileDerived category. A cognate that contains the intra-entity relationship EntityMirrorRemain of a document individual attribute is a cognate of the intra-entity relationship EntityMirrorRemain of a document individual attribute.

[0214] If the cognitive attributes determined by the basic event include one or more document attributes, the cognitive element determined by the cognitive attributes determined by the basic event can be called the cognitive element of the one or more documents.

[0215] Predefined combinations of multiple types of cognitive relationship information, documents, and addressing data for advanced continuation relationships are used to selectively determine whether to create or maintain the addressing data for documents and advanced continuation relationships. Following the order of creating category cognitive elements for advanced continuation relationships and maintaining category cognitive elements for advanced continuation relationships, the advanced continuation relationships identified by the addressing data are updated based on the addressing data for the advanced continuation relationships. For example:

[0216] The inter-entity relationship EntityMirrorChanged including the individual attributes of the document is predefined, and the addressing data of the determined document and the document mirror EntityMirror is newly created.

[0217] If the document operated by the cognize element already has the addressing data of the document and the EntityMirror, the intra-entity relationship EntityMirrorRemain of the document individual attribute is the EntityMirror maintenance category cognize element.

[0218] If the document operated by the cognitive element has no FileFlowNode addressing data, a new category cognitive element of FileFlowNode is determined based on at least one of the following category cognitive elements: a combination of the predefined document mirror entity derivation relationship FileDerived and the document network transmission relationship DataInMotion.

[0219] If the cognitive element already has the addressing data of the document and FileFlowNode, the FileFlowNode maintenance category cognitive element is determined based on at least one of the following category cognitive elements: the document mirror entity derivation relationship FileDerived category cognitive element, the entity intra-relationship EntityMirrorRemain category cognitive element of the document individual attribute, and the document network transmission relationship DataInMotion category cognitive element.

[0220] Table 1 shows a method for determining addressing data of a document operation event destination document high-level continuation relationship.

[0221] Table 1

[0222] The selectivity may be based on a data-driven model or a knowledge-driven model, such as rules.

[0223] An optional approach is to use the source document's hash value as the FileFlowNode.ID value when the source document's extended metadata isn't associated with a FileFlowNode. The FileFlowNode.ID value can also be determined based on content analysis results between documents, including but not limited to document hashes, sizes, and content analysis results.

[0224] Optionally, before step 110 , it is necessary to obtain basic level events of the user on the client.

[0225] For ease of description, the following first describes in detail the process of how to obtain the basic level events of the user on the client with reference to FIG. 2 .

[0226] FIG2 is a schematic block diagram of a basic-level event acquirer 200. Referring to FIG2 , the basic-level event acquirer 200 may include: a file list filter 210, a file event outputter 220, an application API monitor 230, a clipboard event outputter 240, a network event outputter 250, and a print event outputter 260.

[0227] As an example, a program may be installed in a terminal user's device, and when the terminal user accesses one or more documents, detection may be performed using hooks, drivers, and other technologies.

[0228] The functions of the modules included in the basic-level event acquirer 200 are described in detail below.

[0229] The file list filter 210 is used to automatically filter out many insignificant events such as standard calls to generate system files. For example, a file format whitelist is set, and file formats not in the whitelist will not be sent to the basic level event merging / aggregation stage.

[0230] The file event outputter 220 is used to monitor the reading and writing operations of application files, and can be implemented as a driver layer module and / or an application layer module.

[0231] The application API monitor 230 is used to monitor the application running functions, for example, by Hooking, or monitoring program behavior through plug-ins, SDKs, adapters, or modifications.

[0232] The network event outputter 250 is used to identify network behaviors. For example, Windows can monitor network events through different methods such as Winsock, LSP, TDI, and NDIS Driver.

[0233] In the embodiment of the present application, when a user uses a document through an application and an operating system, one or more basic-level events can be output by monitoring the application and / or the operating system using the basic-level event acquirer 200. The output basic-level events may include, but are not limited to, the following information: subject information, event information, object information, time information, etc.

[0234] The above-mentioned subject information may include, but is not limited to, device attribute information, application attribute information, and user attribute information. Device attribute information may include, but is not limited to, device identification (e.g., hard drive serial number, CPU serial number, MAC address), device type, device metadata, port type, driver type, USB device metadata, etc. Application attribute information may include, but is not limited to, application file path, application name, application version, application MD5 hash value, process window title, process company name, process metadata, process start time, process end time, process owner, etc. User attribute information may include, but is not limited to, user ID, user attributes, etc.

[0235] The above event information may include but is not limited to: event name, API name, etc.

[0236] The aforementioned object information primarily refers to the object manipulated by the event. This object is unstructured data at the application level, with read and write access permissions granted to authorized users. Therefore, the object information may include, but is not limited to, source / target file names, file extensions, file modification times, sizes, paths, network operation source / target addresses, ports, and host names, and file extension attributes. It also includes the event's source and destination addresses, network operation information including the source / target file addresses, source and destination ports, host names, and protocols, as well as the number of bytes sent and received, and TCP / IP push and pop event times and counts.

[0237] The above time information may include information such as year, month, day, time, etc.

[0238] It should be understood that the information obtained for different categories of basic-level events may be the same or different, and the embodiments of the present application do not specifically limit this.

[0239] For example, Table 2 below lists some possible basic-level events, as well as examples and related descriptions of the output methods of the basic-level events.

[0240] Table 2

[0241] Basic-level events: opening the application URL and switching the application (URL).

[0242] In the embodiment of the present application, the client can report the basic level event to the event combiner program, which analyzes and processes the basic level event and obtains cognitive relationship information and cognitive attribute information based on predefined rules, which is represented as a document operation event. Several possible implementation methods are described in detail below.

[0243] In one possible implementation, the event combiner merges multiple base-level events to generate document operation events. For example, when an application opens a file, for example by setting a time threshold, if a sequence of multiple "Read File" base-level events is monitored for the same process, the same executable file, the same thread, and the same file handle, the event merging phase will only count a single "Read File" event.

[0244] Another possible implementation involves an event combiner correlating multiple base-level events to generate document operation events. For example, the event combiner can analyze application behavior, consider subject characteristics, event characteristics, and object characteristics included in base-level events, and selectively aggregate eligible base-level events to generate document operation events.

[0245] The subject features mentioned above may include: user information, application information, device information, and other subject information when the basic-level event occurs. As an example, in an embodiment of the present application, subject and subject attribute features can be used to associate different basic-level events together, including basic event association based on interface features. When it comes to user network operations, the application window name, browser tab name, application title bar, etc. can be used for basic event association.

[0246] These event features may include: base-level events and base-level event attribute features, specifically: event category, event count, read byte count, written byte count, event start time, event end time, and source location information. Depending on the type of base-level event, additional information such as URI, UNC, and URL may also be included.

[0247] The above-mentioned object characteristics may include: files and file attributes of basic-level events.

[0248] In an embodiment of the present application, when the event combiner detects that the basic-level events are combined in a predetermined pattern, for example, the time sequence of the occurrence of the basic-level events conforms to a predefined time feature, it is considered to be a specific document operation event. The predefined time feature refers to a time sequence composed of multiple basic-level event triggering times. The event combiner can also obtain the monitoring results of a sequence of data asset access events. The monitoring results are a sequence of data asset access events and the corresponding subject and object information, specifically including: source location, destination location, source document, destination document, calling program ID, executable program name, start time, end time, login or logout user operations, time and user identity identification, device type and other information.

[0249] The following examples illustrate different document operation events obtained based on basic level events.

[0250] In Example 1, the document operation event is "File Created." The basic events that make up this document operation event include the following basic-level events: "Open File" and "Write File." The combined characteristics of the basic events of this document operation event can be understood as right-click Create and Application Create.

[0251] Next, we analyze the creation of new processes within applications. For example, in complex applications like Office and WPS, the main process is solely responsible for interface interaction and rendering, while each session logic within the application is implemented by one or several separate sub-processes. It's necessary to aggregate the basic events of each sub-process, sort them by occurrence timestamp and target object, and then process them uniformly.

[0252] Scenario 1: After child process A reports "Open File" for the c:\test.docx document, it performs several "Read File" operations. Child process B then performs a "Write File" basic event on the c:\test.docx document. Because the predefined time signature for "File Edited" is met and the editing window associated with the c:\test.docx document exists, a "File Edited" file content-related event is generated.

[0253] Scenario 2: After child process A reports "Open File" for the c:\test.docx document, it performs several "Read File" operations. Process B "Opens File" for the new document c:\test1.docx and performs the "Write File" basic event. Furthermore, the document c:\test.docx is unlocked, the document c:\test1.docx is locked, and the editing window associated with the document c:\test.docx is associated with the new document c:\test1.docx. After combined analysis, a "Save as New File" event for the document c:\test.docx to c:\test1.docx is generated in the software.

[0254] Scenario 3: After child process A reports "Open File" for the c:\test.docx document and then performs several "Read File" operations, if process B "Opens File" for the c:\test1.docx document and only performs the "Write File" basic event, and a new editing window appears associated with the c:\test1.docx document, and the c:\test.docx document is not unlocked, and the associated window still exists, then after combined analysis, a new event for the c:\test1.docx document will be generated in the software.

[0255] The process of correlating basic level events involves the following implementations:

[0256] Method 1: When the application opens a file, it caches the "Open File" path and window title, selects "Write File" as the target file for Save As, obtains file information through the current window title, and finds the corresponding file in the cache as the source document for Save As.

[0257] Method 2: Get the current window title.

[0258] It should be understood that different application titles may have different formats. Some titles include the full file path, some include the file name, and some include only the short file name without the suffix. For example, if the full file path is included, the original file can be obtained directly from the title. For another example, if the file name is included, the cache is searched, the file name is matched, and the corresponding original file is obtained.

[0259] For example, the formats of titles are mainly as follows:

[0260] The first type: the file name comes first, such as test.docx office; the second type: the file short name comes first, such as test office;

[0261] The third type: the full file path comes first, such as C:\test.docx office; the fourth type: the file name is enclosed in [], such as [test.docx]office; the fifth type: the full file name is enclosed in [], such as: [C:\test.docx]office.

[0262] Optionally, there are other types, and you only need to provide the corresponding file name information for these formats.

[0263] Optionally, you can filter out files with unnecessary suffixes and files in temporary and specific directories, so that the files read by the application are cached by the client.

[0264] Both WPS and Office applications support Microsoft's document COM component secondary development technology. They can obtain the document's COM component object through the document window handle, and thus directly extract the full document path corresponding to the document window. This can effectively solve the problem of confusion that may occur when documents with the same name but different paths are opened at the same time through title bar matching.

[0265] The following describes how to obtain the source and destination documents for this type of document operation. The source path for file migration operations between different disk paths can be obtained in a variety of ways. Whenever a file is opened and read, it can be cached. Most applications occupy files while operating on them, so this feature can improve how files read by applications are cached.

[0266] Whenever a file is opened and read, the monitoring program puts the file into the cache; when the file is closed, the monitoring program deletes the file from the cache. In this way, the cache only stores the files currently being operated by the program, and the cache size is always within a controllable range.

[0267] Some applications close the current file before saving or saving as a file. This can cause the file to be cached by the application. To overcome this problem, you can add a special cache that always saves the last N files closed. This method has the advantages of the application-occupied file cache method without this particular problem.

[0268] Some special applications generate a temporary file when opening a file. This file is then opened and the original file is closed simultaneously. In this case, the original file is no longer in the cache. In these cases, the original file can be retrieved using a special image of the temporary file. For example, a special $ character is always added to the beginning of the document. When the client monitor reads the cache and finds the first character to be a $, the "File Edit" function uses this image to retrieve the original file.

[0269] Example 2, the document operation event is "save as a new file". The basic events that make up this document operation event include the following basic-level events: "read file", "read and write file", and "write file". The scope of basic events associated with this document operation event includes: basic event association based on subject attributes and association based on interface features. The combined features of the basic events of this document operation event can be understood as an application "reading a file" or "reading and writing" a file, and then "writing a file" to a new file. Therefore, if the occurrence sequence of the basic-level events meets the above features, a document operation event of "save as a new file" can be obtained.

[0270] The following describes how to obtain the source document entity and destination document under this document operation event.

[0271] Method 1: Basic events based on interface feature association.

[0272] The monitoring program caches the file opened by the basic event "Open File" and the associated window of the file; some applications' "Save as New File" operation requires a write file operation. The monitoring program obtains the file written when the "Write File" basic event operation occurs as the target file of "Save as New File" (Note: files in the temporary directory are not save file operations, files with specific suffixes belong to save file operations, and file names with specific formats belong to save file operations). The current associated file path is obtained through windows, plug-ins, etc. If the window was originally associated with the object (file) that was previously operated by the process's "Read File" basic level event or "Read and Write" basic level event, and is now associated with the target file of the "Write File" basic event, the source document entity of the "Save as New File" operation event can be determined.

[0273] Method 2: Use the plug-in program to accurately obtain the source file and destination file.

[0274] Office supports COM plug-ins, the interface is Extensibility, implement this interface, implement OnAddInsUpdate, OnBeginShutdown, OnConnection, OnDisconnection, OnStartupComplete methods, you can get document open, close, save, "Save As" and other events.

[0275] Adobe supports DLL plug-ins. The DLL exports a specified function PlugInMain to initialize the context and receive events such as document opening, document closing, document switching, and document "Content Save As".

[0276] When editing a document, editing software such as Office / WPS will automatically lock the document being edited, meaning that other third-party software can no longer open and access the document. (Supplement: In some special cases, the document may not be locked during the entire reading or editing process. In this case, it is necessary to perform a logical supplementary judgment on the document lock by checking whether the document-related window exists, rather than relying solely on the file lock status.) Therefore, when a Save As event occurs, the target software will behave by unlocking the current document, locking the new file, and performing write access to the new file. Therefore, combining the basic events of this sequence will allow the source and destination documents of the Save As event to be jointly inferred.

[0277] Method 3: Use collaborative data flow between applications to obtain source files and destination files.

[0278] Multiple processes collaborate to complete the Save As operation. Sometimes, operations such as reading the source file, converting the format, creating the destination file, and writing data are all completed collaboratively between different processes and are not limited to one process. For example, WPS provides several auxiliary editing tools to complete batch conversions between various document formats (mainly Office documents to PDF or image formats). In this case, simply using the methods mentioned in Method 1 or Method 2 cannot meet the event capture requirements. Therefore, it is necessary to first intercept the initial document path passed in at startup from the gadget, and combine it with Method 1 to record the captured basic read and write events and timestamps of the source file and destination file in multiple processes. Finally, when the basic event of the destination file is captured, use the timestamp to query the basic event information recorded in other processes to integrate the source file path (the purpose of using the timestamp is to filter out those read events that occurred earlier) to complete the collection and reporting of SaveAs event information.

[0279] Example 3, the document operation event is "file is copied". The basic events composed of this document operation event include the following basic-level events: "read file", "write file", "read and write file" and "copy file". The basic event scope associated with this document operation event is the basic event association based on the subject attributes, source document: the document operated by the "read file" basic event. Destination document: the document operated by the "write file" basic event or the file written by the "copy file" basic event. The combined characteristics of the basic events of this document operation event can be understood as the same thread, and the alternating file reading and file writing processes can be monitored. File reading has the same file handle, and file writing has the same file handle.

[0280] Example 4, the document operation event is "the file is burned". The basic events that make up this document operation event include the following basic-level events: "open file", "read file", "write file" and "rename file". The combined characteristics of the basic events of this document operation event can be understood as, for example, some burning files can be composed of basic-level events such as "open file", "read file", "write file", etc., and the subject of the basic event is the burning application. The objects of the basic events involved may involve temporary files. Basically, by observing the behavioral characteristics of various commonly used burning programs, a logic of a predefined basic event sequence applicable to almost all commonly used burning programs can be formed.

[0281] Example 5: The document operation event is "file compressed (zip)." The basic events that make up this document operation event include the following basic-level events: "open file," "read file," "write file," and "rename file." The combined characteristics of the basic events of this document operation event can be understood as observing the behavioral characteristics of various compression programs. A "file compressed" event is composed of basic-level events such as "open file," "read file," "write file," "rename file," and "move file." Through the association of the subject-object environment time characteristics of these basic events, the occurrence of these basic events conforms to a predefined time sequence.

[0282] The following is an analysis of the source and destination document characteristics of mainstream compression software on the market:

[0283] Winrar: First "open file" destination document (compression type document such as rar\zip), then "open file", "read file", "file is compressed" doc\pdf and other user files (as the source document of "file is compressed"), then "write file" rar\zip and other compression type documents (as the destination document of "file is compressed").

[0284] haozi, 360zip: First "open file" temporary file (*.tmp type) in the destination path, then "open file" and "read file" compressed doc\pdf and other user files (the source document of "file compressed"), then "write file" temporary file, and then "rename" the temporary file to rar\zip and other compression type documents (the destination document of "file compressed").

[0285] Winzip: "Open file" and "read file" doc\pdf and other user files (as the source document of "file is compressed"), then "write file" temporary file (random file, *.tmp file), then "rename" the temporary file to the destination document of the temporary path (rar\zip and other compression type documents), then "move file" such as zip from the temporary path to the target path specified by the user (as the destination document of "file is compressed").

[0286] When analyzing program behavior, you can sometimes use specific parameters to judge, such as GENERIC_WRITE or FILE_WRITE_DATA or FILE_WRITE_ATTRIBUTES parameters, which can be considered as the compressed destination file name path.

[0287] Example 6: The document operation event is "file is unzipped." The basic events that make up this document operation event include the following base-level events: "open file," "read file," "write file," and "rename file." The combined characteristics of the base events of this document operation event can be understood as observing the behavioral characteristics of various compression programs. A "file is unzipped" event is composed of base-level events such as "open file," "read file," "write file," "rename file," and "move file." These base events are associated with the subject, object, and environment time characteristics of these base events, and the occurrence of these base events conforms to a predefined time sequence.

[0288] For example, the monitoring of winrar first "opens the file" and "reads the file" rar\zip and other compression type documents (as the source file of "file being decompressed"), then "opens the file" doc\pdf and other user file types (as the destination file of "file being decompressed"), and then "writes the file" to the destination document.

[0289] When analyzing program behavior, you can sometimes use specific parameters to make judgments. For example, when decompressing a generated file, the program will set file parameters.

[0290] In Example 7, the document operation event is "file archived." Similarly, events such as file archiving require observation of application behavior. Characteristics include the application being archived and the event characteristics of the archive operation. For example, it consists of basic-level events such as "open file," "read file," "write file," and "rename file," and their sequences. The objects involved include temporary files and user files.

[0291] In Example 8, the document operation event is "File uploaded." This event is represented by "JustUpload" in the policy. The basic events that make up this document operation event include the following basic-level event: "Select File." When a web application performs an operation such as "Select File," a redirect is performed, allowing the web application to access the newly created file stored in a temporary path. This operation event does not require network information.

[0292] Example 9, the document operation event is "the file is uploaded and the network information is obtained", which is represented by Upload in the policy. The basic events that make up this document operation event include the following basic-level events: "network connection", "select file", and "window name changes". The scope of basic events associated with this document operation event includes basic event associations based on subject attributes and associations based on interface features. The source document is the document operated by the "select file" basic event, the destination document is the redirected file, the source location is the document path operated by the "select file" basic event, and the destination location is the network address connected by the "network connection" basic event. The combined features of the basic events of this document operation event include a combination of basic event sequences such as "network connection", "window name changes", and "select file".

[0293] Step 1: Association between the “network connection” basic event and the application body (interface features).

[0294] The embodiment of the present application sets a time interval (for example, 1 second) and utilizes causal relationships and time intervals to associate network IP addresses with application interface features. The "network connection" basic event occurs first, with the IP address connected and the content transmitted being the cause. The result is a change in the application interface content (browser tab / application window title), i.e., the "window name changed" basic event occurs.

[0295] For example, if the Chrome browser tab name changes at time p, then the one or more new IP connections established by Chrome.exe within the previous P-1 seconds are the cause of the change in the current window's tab name and are the possible IP addresses corresponding to the process / thread. The tab name can be associated with the corresponding IP address and saved. For example, if the Chrome browser tab was updated at 9:23:15, or 09:231500 seconds, and a predefined time interval is added forward, the start time is 9:23:14, or 09:231400 seconds. The URL can be used as an auxiliary identification method for the IP address.

[0296] One scenario involves a browser determining the internal IP address. When a browser or application accesses an internal server, the browser tab changes at time K. One second before time K, the browser establishes new connections to multiple IP addresses. If one of these IP addresses is an internal IP address, then that IP address is determined to be the server IP address accessed by the browser tab. This is because internal servers operate purely for advertising and other services, with only one server communicating with the client. If no new sessions are established with the internal IP address within the time interval, the application tab or window is likely connecting to an external network. It's worth noting that this method accurately determines the internal IP address to which the application is connecting. This is because enterprises have a small number of internal servers, operate purely for advertising and other services, and the client directly transmits content after connecting to the server. When determining external IP addresses, further corrections and analysis are required based on the type of service.

[0297] Another scenario involves the browser determining the public IP address. When a browser or application accesses an intranet server, the browser tab changes at time K. One second before time K, the browser established new connections to multiple IP addresses. If all of these IP addresses are public IP addresses, the browser can optionally analyze the application's characteristics and set different judgment models based on the application type to identify the most likely server IP address that the webpage primarily connects to. 1) For example, after analysis, some applications can determine the server IP address based on the most recent time interval. For example, the IP address with the closest time interval between the connection establishment and the browser tab update can be used as the tab's IP address. 2) For example, after analyzing the downlink and uplink data counts within a time interval, some applications can determine the server's most likely public IP address based on the IP address with the largest number of downlink packets. 3) For example, some browsers can obtain the URL to further assist in determining the IP address corresponding to the tab / window name. During application use, the IP address can still be modified based on the connection to identify the most likely IP address among multiple IP addresses.

[0298] One auxiliary identification method for obtaining the URL of the browser's current page is to use the AccessibleObjectFromWindow system call to obtain the IAccessible interface object associated with the browser window page, which is then used to traverse the content in the address bar of the current page, which is the URL of the current page; then, the server domain name in the URL is parsed using the DNS protocol to obtain the server IP address of the current page. However, some browsers do not directly support the IAccessible interface to obtain the actual page element content by default, and support must be enabled through startup parameters (for example, Chrome requires the startup command line parameter --force-renderer-accessibility to enable support for this feature).

[0299] Another scenario is an intranet client-server application (CS application). This type of client typically connects to a single server. The "Network Connection" basic event is associated with the "Window Name Change" basic event within a time interval. This allows the IP address to be associated with the application's main characteristics (tabs).

[0300] Step 2: Association between the application body (interface features) and the "select file" basic event.

[0301] When the "Network Connection" basic event is generated, the correspondence between the application title name, tab name, window name and IP is cached as the IP address of the application window. When the "File Selection" basic event occurs, the event attributes such as process and current window can be used to associate it with the current window of the application where the selected file was uploaded.

[0302] Therefore, the IP address associated with the application window can be used as the destination address (IP address). The document operated by the "Select File" basic event is the source document. The redirected file is used as the destination document.

[0303] In Example 10, the document operation event is "File downloaded and network information obtained." The base events represented by "download" in the policy include the following base-level events: "Network connection," "Copy file," "Move file," "Rename file," "Open file," "Read file," and "Write file." The download source is the network address connected to the "Network connection" base event, the source document is the file that meets the predefined logic of the file downloading base event, and the destination document is the document after the download is complete and the document's cognition metadata has been updated.

[0304] There are many ways to determine the target file to be downloaded. The embodiments of the present application do not specifically limit this. As an example, the file operation sequence can be composed of basic events such as "copy file", "move file", "rename file", etc., which can be used to determine the destination document to be downloaded. As another example, by monitoring basic events such as "open file", "rename file", "move file", etc., and analyzing the characteristics of temporary folders and temporary files of basic event operations, it can also be used to determine the destination document of <file is downloaded and network information is obtained>. It should be understood that some application download processes will only generate one type of temporary file, and some application download processes will generate two or more types of temporary files in sequence.

[0305] As an example, the following describes the file operation sequence of mainstream browsers.

[0306] Chrome and Edge browsers: Downloads generate multiple temporary files, so the download logic flow goes through multiple "Open File" and "Rename File" events for these temporary files. These events include "Open File" (creating a temporary *.tmp file), "Rename File" (renaming *.tmp to *.crdownload), and "Rename File" (renaming the temporary crdownload file to the download target file).

[0307] 360 Browser: "Network Connection", "Open File" (create a temporary file *.dl), "Rename File" (rename the temporary file to the target file downloaded by the user, for example, if *.dl is renamed to 123.doc, then 123.doc is the downloaded destination document.

[0308] IE browser: "Network Connections", "Open File" (create a temporary file *.partial), "Rename File" (rename the temporary file to the target file downloaded by the user, for example, rename *.partial to 344.xls, then 344.xls is the downloaded destination document).

[0309] QQ Browser: "Network Connection", "Open File" (create a temporary file *.qbl), "Rename File" (rename the temporary file to the target file *.qbl downloaded by the user to kk.pdf, then kk.pdf is the downloaded destination document).

[0310] Opera browser: "Network Connections", "Open File" (create a temporary file *.opdownload), "Rename File" (rename the temporary file *.opdownload to the target file downloaded by the user).

[0311] Maxthon browser: "Network Connection", "Open File" (create temporary file *.crdownload), "Rename File" (rename the temporary file to the target file *.crdownload downloaded by the user).

[0312] Sogou browser: "Network Connection", "Open File" (create a temporary file *.sgdownload), "Rename File" (rename the temporary file to the download target file).

[0313] The combined features of the basic events contained in the above-mentioned "file downloaded and network information obtained" document operation event are as follows:

[0314] The first type of application feature 1: "Network connection", "Open file", "Rename file" event sequence.

[0315] The second type of application feature 2: "Network connection", "Open file", "Rename file", "Open file", "Rename file" event sequence.

[0316] The following describes the combination of basic events.

[0317] Step 1: Associate the "network connection" basic event with the application subject.

[0318] Similar to the "network connection" in "File is uploaded and network information is obtained", the interface characteristics such as application window, title, and label (application window update event or tab name) are associated with the IP address.

[0319] Taking advantage of the fact that browsers generate temporary files immediately when users click the download button, we can associate the creation time of the temporary file in the "Open File" function with the current application window. This allows us to associate the IP address / URL obtained through the "Network Connection" function with the corresponding application (interface features) and the temporary file created by the "Open File" function.

[0320] Some applications can use other attributes of the subject, such as the process ID, for association. For example, the process ID of a "network connection" is also the process ID of a temporary file created by "opening a file." This allows for association and combination of different basic events.

[0321] Step 2: Association between the application program body (interface features) and the basic event sequence of file operations.

[0322] When a user downloads a file while browsing a web page, a new temporary file is created in the browser's temporary folder. This can be understood as a new temporary file, which is an operation performed by the user while browsing the current tab.

[0323] The destination document determination method of "the file is downloaded and network information is obtained" mainly analyzes the characteristics of the application download: when the browser and network application download the file, a specific download logic flow will be executed.

[0324] Through the "network connection" basic event, you can first associate the IP with the application window name of the "window name changed" cache. If for some specific applications, the window has a "new file" basic event for a temporary file in a predefined temporary folder, and then the temporary file has a "rename file" to a user-common document format such as doc, then a "file downloaded and network information obtained" event is generated, where the source address is the IP address associated with the main window, and the destination document of the "rename file" basic event is used as the destination document of the operation event.

[0325] The following takes the Chrome browser as an example to briefly describe the analysis and processing process (other browsers are similar).

[0326] The basic event sequence corresponding to Chrome browser <File is downloaded and network information is obtained> is: "Network connection", "Open file" (*.tmp temporary file), "Open file" (*.crdownload temporary file), "Rename file" (rename tmp to crdownload), "Open file", "Rename file" (rename the crdownload temporary file to the download target file).

[0327] For example, a user has multiple tabs open in the Chrome browser: tab1: Yahoo, tab2: Gmail, tab3: Gmail Inbox, and tab4: Facebook. The "Network Connections" function retrieves the IP address of tab1: 34.54.98.4; tab2: 56.48.32.12; tab3: 56.48.32.12; and tab4: 86.12.144.23. Because each tab accesses a different website, when you click the Download File button on the application interface, the webpage or application displayed in the current window is the source of the downloaded file, i.e., the source IP address. The current window's time is recorded. For example, the current window's time for tab3 is 9:45:22 to 9:53:34, and the current window's time for tab2 is 9:43:01 to 9:55:21.

[0328] Monitoring revealed a predefined sequence of basic events ("network connection," "open file," "rename file," "open file," "rename file") in the Chrome browser, matching the predefined sequence of basic events for "file downloaded and network information obtained." The "open file" event for the temporary *.tmp file occurred at 9:47:43. The current Chrome browser window corresponds to tab 3, indicating that the IP address 56.48.32.12 corresponding to tab 3 is the source IP address for "file downloaded and network information obtained." The destination file, demodownload.docx, generated by the second "rename file" event, is also known as the destination file.

[0329] The following is an analysis of the file operation sequence (CS client software).

[0330] Some client-server applications first create a temporary file in the user-selected download destination folder, then repeatedly write to the temporary file, and finally rename the temporary file to the target file. Some CS applications create a temporary file in a temporary folder, rename the file to the target file format, and then move the file to the final download destination.

[0331] In Example 11, the document operation event is "File downloaded (no IP address determination required)" and the policy uses "JustDownload." This document operation event comprises the following basic events: "Copy File," "Move File," "Rename File," "Open File," "Read File," and "Write File." Determining the download destination file is similar to determining the destination file for "File downloaded and network information obtained."

[0332] Example 12: The document operation event is "file moved". The basic events of this document operation event include the following basic-level event: "Move file".

[0333] Example 13: The document operation event is "file renamed". The basic events of this document operation event include the following basic-level event: "file renamed".

[0334] In Example 14, the document operation event is "File Edited." The basic events comprising this document operation event include the following base-level events: "Read File," "Write File," "Read and Write File," "Copy File," and "Move File." The scope of basic events associated with this document operation event includes base event associations based on subject attributes, namely, the document operated on by "Write File" and the document operated on by "Read and Write File." The combined features of the basic events contained in this document operation event include Features 1 and 2, where Feature 1 satisfies any of the following conditions: the same thread "reads" and "writes" the same file handle; the same process ID and the same file handle, and the file hash changes after the first read and the last write. Feature 2 includes: Some applications generate temporary files when opening a file, and the edited content is saved in memory or a temporary file. When the application saves the edited content, it first writes the content to the temporary file and then saves the contents of the temporary file to the destination document through base-level event operations such as "Copy File" and "Move File."

[0335] Example 15, the document operation event is "file is read-only". The basic events that make up this document operation event include the following basic-level events: "open file", "read file" and "close file". The combined characteristics of the basic events contained in this document operation event include: 1. Any process started by the same program performs "open file" on the same file; 2. Subsequently, the program performs a "read file" operation, then the file records the time of the first "read file" event as the start time of "file is read-only"; 3. After the file is marked as read-only and valid, after the "file close" event occurs (the open count of all files of different processes of the same program is cleared to zero), when the file is not occupied (that is, the "unlock file" event occurs), the "file is read-only" event ends.

[0336] It should be noted that after a "read file" occurs, it is necessary to determine whether the file is occupied. If the occupation time exceeds a certain length of time (such as 2 seconds), the "file read-only" status is confirmed to be valid. Otherwise, the file is only temporarily accessed by the application (such as adding it to the recent access list, establishing an access index, etc.). The determination that the file is occupied includes determining that the file cannot be opened exclusively by CreateFile, that is, it is determined to be occupied. For some special applications, the file is not locked during editing. Instead, it is necessary to find the editing window corresponding to the file and the application interface. If the corresponding window exists, the file is determined to be occupied; if the corresponding window does not exist, the file is determined to be unlocked.

[0337] Example 16, the document operation event is "content is pasted to the clipboard". The basic events that make up this document operation event include the following basic-level events: "clipboard copies content" and "clipboard pastes content". The scope of basic events associated with this document operation event is the basic event association based on subject attributes. The combined features of the basic events contained in this document operation event include: "clipboard copies content" and "clipboard pastes content". When a user reads or edits a document, the action of copying the document content to the clipboard occurs, and the feature code of the copy sequence content is recorded, and the feature code, timestamp and document path are cached together in the shared cache area of ​​the local computer; when a paste event occurs, the content is taken from the clipboard, and after generating a feature code, its source is retrieved from the shared cache area. If the most recent copy matching source record is found, the source document of the paste action is found, and a record of the document's "content is pasted to the clipboard" event action can be formed.

[0338] The following examples illustrate the document mirror entity derivation relationship (FileDerive) in combination with different document operation events.

[0339] In Example 1, the document operation event is "File Created." Because this document operation event represents the inter-entity relationship of individual document attributes, a new document ID needs to be created and associated with the destination file. The document mirror entity derivative relationship (FileDerive) can be expressed as: (0: destination file ID).

[0340] Example 2, the document operation event is "save as a new file". Since this document operation event expresses the entity relationship of the individual attributes of the document, it is necessary to create a new document ID and associate it with the destination file. The document mirror entity derivation relationship can be expressed as: (ID of the source file: ID of the destination file).

[0341] Example 3, the document operation event is "file is copied". Since the document operation event expresses the entity relationship of the individual attributes of the document, a new document ID is created and associated with the destination file. The document mirror entity derivative relationship can be expressed as: (ID of the source file: ID of the destination file).

[0342] Example 4, the document operation event is "file is burned". Since the document operation event expresses the entity relationship of the individual attributes of the document, a new document ID is created to associate with the destination file. The document mirror entity derivation relationship can be expressed as: (ID of the source file: ID of the destination file).

[0343] In Example 5, the document operation event is "file compressed (zip)." Since this document operation event expresses an inter-entity relationship between individual document attributes, a new document ID is created and associated with the destination file. This document mirror entity derivation relationship can be expressed as: (source file ID: destination file ID). The source file ID is the identifying attribute of the source file on disk that the user intends to compress, and the destination file ID is the new identifying attribute of the compressed file actually obtained by the compression program.

[0344] In Example 6, the document operation event is "file decompressed (unzip)." Because this document operation event expresses an inter-entity relationship between individual document attributes, a new document ID is created and associated with the destination file. This document mirror entity derivation relationship can be expressed as: (source file ID: destination file ID). The source file ID is the ID stored in the user file (e.g., .doc) generated by the compression program after decompression, and the destination file ID is the ID of the file that replaces the original decompressed file.

[0345] The following uses different document operation events to illustrate how a document is transmitted over the network using DataInMotion.

[0346] In Example 1, the document operation event is "File Uploaded." Because this document operation event expresses the inter-entity relationship between individual document attributes, a new document ID is created and associated with the destination file. The relationship category for this document being transferred over the network can be represented as: (Source File ID: Destination File ID). The source file ID is the ID of the document stored in the "Select File" basic event, and the destination file ID is the ID of the redirected file. This event does not require determining the IP address; what is transferred is a copy of the disk-stored file.

[0347] In Example 2, the document operation event is "File uploaded and network information obtained." Because this document operation event expresses the inter-entity relationship between individual document attributes, a new document ID is created and associated with the destination file. The document network transmission relationship can be expressed as: (Source file ID: Destination file ID). The source file ID is the ID of the document stored by the user to upload, and the destination file ID is the ID stored in the redirected file.

[0348] In Example 3, the document operation events are "File downloaded and network information obtained" and "File downloaded (no IP address determination required)." Because these document operation events express the inter-entity relationship between individual document attributes, a new document ID is created and associated with the destination file. The document network transmission relationship can be expressed as: (Source file ID: Destination file ID). The source file ID is the ID of the document after the file download is complete, and the destination file ID is the ID of the new file that is replaced and stored on disk after the network application download operation is completed.

[0349] For example, if the source document ID in a "file is downloaded and network information is obtained" operation event is the ID of the destination document in another event "file is uploaded", then the two documents belong to the same FileFlowNode high-level continuation relationship.

[0350] The subject information mentioned above may correspond to the subject information mentioned above, including but not limited to: device attribute information, application attribute information, and user attribute information. Please refer to the description above for details, which will not be repeated here.

[0351] The above folder information can be determined by the path information of the target document, for example, if FilePath=c:\123\44.doc, then FolderPath=c:\123\.

[0352] After obtaining the cognitive relationship information and cognitive attribute information, the addressing data of the high-level continuation relationship is associated with the document.

[0353] There are several options for associating the addressing data of a high-level continuation relationship with a document:

[0354] 1) The addressing data of the advanced continuation relationship is directly stored in the extended attributes of the document, and the relationship between the document attributes and the addressing data of the advanced continuation relationship is established. For example, the ID of FileFlowNode.ActionBusiness is stored in the extended attributes of the document.

[0355] 2) Create an individual identifier and store it in the document's extended attributes, establishing a relationship between the original document's individual attributes and the individual identifier in the extended attributes, and establishing a relationship between the individual identifier stored in the extended attributes and the addressing data of the higher-level continuation relationship. For example, create a FileID and establish a relationship between the FileID and the addressing data of the higher-level continuation relationship, including storing the relationship between the document mirror EntitiMirror.ActionBusiness and the addressing data of the higher-level continuation relationship in the database, and the EntityMirror.ActionBusiness identified by the individual identifier FileID stored in the extended attributes; the document entity flow relationship FileFlowNode.ActionBusiness is associated with or stores the FileID or EntityMirror.

[0356] 3) Based on the document location, select a storage method for the addressing data of the higher-level continuation relationship corresponding to the document. When the epistemology indicates that the destination document is at risk of losing the addressing data of the higher-level continuation relationship of the source document, ensure that the addressing data of the higher-level continuation relationship of the destination document is not lost.

[0357] That is, the combination of basic events, in addition to updating high-level continuation relationships, can maintain the addressing relationship between the original attributes of the document and the high-level continuation relationship based on the decision center, so that the addressing data of the document and high-level continuation relationships such as FileFlowNode can flow with the flow of the document without separation. For example, based on the document storage location, the storage method of the addressing data of the high-level continuation relationship is determined; a combination of multiple storage methods is used to store the addressing data of the high-level continuation relationship. For example, the relationship storage method of the original attributes of the document - the addressing data of the high-level continuation relationship can realize the collection of document cognition that collaborates on multiple devices within the enterprise.

[0358] In some embodiments, during the process of determining a cognate, at least one of the following methods may be selected to store extended attributes of the document operated on by the cognate based on the storage location of the document operated on by the cognate, where the extended attributes include addressing data of a high-level continuation relationship:

[0359] Method 1: Embed the extended attributes into the extensible attributes in the file format;

[0360] Method 2: used to modify the document content storage extended attributes;

[0361] Method 3: Encrypt and encapsulate the extended attributes and the main text file into one file;

[0362] Method 4: Store extended attributes in the file system's extensible attributes section;

[0363] Method 5: Store extended attributes in a predefined database or file.

[0364] For ease of understanding, the above-mentioned methods 1 to 5 are introduced in more detail below in the form of examples and will not be described in detail here.

[0365] The module (or component) storing the addressing data of the advanced continuation relationship in the document extended attribute is described in detail below.

[0366] As an example, the modules for storing addressing data of high-level continuation relationships may include but are not limited to: ExtensiveTag, InvisibilityTag, UserEditTag, DatabaseTag, EncryptedTag, FileSystemTag. The functions of the above modules are described in detail below.

[0367] ExtensiveTag represents an extensible attribute that embeds addressing data of advanced continuation relationships into the file format. The specific functions of ExtensiveTag include but are not limited to the following:

[0368] 1) Use metadata storage specifications developed by international organizations or corporate alliances, such as documents that support ODF, OOXML, XMP, and UOF standard formats. Addressing data for advanced continuation relationships can be stored in predefined (custom attribute) locations in the document format. For example, OOXML supports Microsoft Office documents such as docx and xlsx, and the ODF format supports odt, ods, fods, odp, fodp, odg, fodg, and odf formats. For example, the XMP standard format (Extensible Metadata Platform) can embed addressing data for advanced continuation relationships into file formats such as pdf, jpg, DNG, GIF, JPEG, PNG, TIFF, MP3, MPEG-2, MPEG-4, SWF, HTML, and XML.

[0369] 2) Application program interface API storage provided by application manufacturers, such as WPS API, can store metadata in the WPS series file formats.

[0370] 3) Store the addressing data of the high-level continuation relationship in a public file format that can store metadata, such as PDF.

[0371] 4) The addressing data of the high-level continuation relationship is stored in the location of the user attribute in the file format. For example, the author attribute can be stored in the docx file format.

[0372] InvisibilityTag is used to modify content and store metadata. Specifically, depending on the file format of the carrier, such as audio or video data, it can be used to store addressing data with high-level continuity relationships in audio or video images through LSB substitution steganography, MLSB substitution steganography, random modulation steganography, etc., making the modifications imperceptible to the human visual and auditory systems.

[0373] UserEditTag is used to modify content and store the addressing data of high-level continuation relationships. Specifically, for many format documents, users can use this component to store the addressing data of high-level continuation relationships in the document's notes attribute.

[0374] DatabaseTag indicates that the addressing data of the high-level continuation relationship is stored in the database. In the embodiment of the present application, the addressing data of the high-level continuation relationship can be combined and associated and added to the DB as a DB tuple.

[0375] EncryptedTag indicates that the addressing data and the main file of the high-level continuation relationship are encrypted and encapsulated into a single file. Optionally, EncryptedTag is also responsible for decrypting the encrypted and encapsulated metadata. In one optional implementation, the encrypted and encapsulated metadata record format can include three parts: a header, a metadata portion, and the encrypted data. The header can include an identifier, such as a key identifier or a document identifier.

[0376] FileSystemTag represents addressing data for storing high-level continuation relationships within the file system's extended attributes. For example, Windows NTFS files have a corresponding ADS data stream (Alternate Data Streams). The NTFS file system supports variable data streams of unlimited length. FAT, HPFS, NTFS, ext4, and JFS also support file system extended attributes.

[0377] In some embodiments, document operations may result in the loss of metadata associated with the document, posing a risk of failing to accumulate the source document's high-level continuation relationship cognitive attributes. For example, when a user performs a "Save as New File" operation, such as saving 456.doc as 999.pdf, the newly generated file loses the source document's high-level continuation relationship addressing data. Therefore, after the Save as operation, the source document's high-level continuation relationship addressing data can be stored and associated with the destination document. This overcomes the issue of different document operations causing a failure to continuously accumulate the document's high-level continuation relationship cognitive attributes. Furthermore, the "Save as New File" cognitive relationship information and cognitive attribute information are transmitted to the server, enabling the continuous accumulation of the document's cognitive metadata associated with the document operation and the document's corresponding high-level continuation relationship cognitive attributes.

[0378] In one possible implementation, a storage method for addressing data of a high-level continuation relationship may be selected based on the attributes of the document's location, and the storage method for addressing data of the high-level continuation relationship may be automatically updated in response to changes in the document's location.

[0379] For example, Table 3 below lists the storage method of the document location and the document-related metadata.

[0380] Table 3

[0381] For example, if the document is located on a hard disk within the enterprise network and can be monitored by a monitoring program, you can choose to store the addressing data of the advanced continuation relationship in three locations at the same time: 1 <filesystemtag>File system extended attributes, 2 <extensivetag>The extended attributes and 3DatabaseTag database in the file format allow the document to be identified even if the content is encrypted, as the addressing data of the advanced continuation relationship is stored in the DB or in the file system extended attributes.

[0382] Sometimes, when a file is edited and saved, for example, when certain applications open, read, or write a file, their logic is to first delete the file and store a new file with the same name and path. When a file is opened (e.g., when creating a file), the addressing data for the file's high-level continuation relationship is saved. When a file with the same name is deleted or created, the original high-level continuation relationship addressing data can be restored to the file format's extended attributes, the file system's extended attributes.

[0383] An alternative approach is to upload the document to a server that cannot be monitored by monitoring programs (inside the corporate network), and to store metadata such as addressing data of advanced continuation relationships in addition to predefined locations in the storage document format ( <extensivetag>), the addressing data of the high-level continuation relationship is still embedded in the file and is still associated with the document. And the document is in plain text and can be recognized by software such as ERP and CRM. It can also be encoded to <useredittag> , <invisibilitytag>When a "file is downloaded and network information is obtained" event occurs, the client obtains the destination document and waits for the destination document's "file to be closed". It can then verify whether the addressing data of the stored advanced continuation relationship has been modified. If not, it can be deleted. <useredittag>This addresses the issue of a few non-compliant application vendors on the internet deleting user-defined metadata stored in the file format. Most compliant vendors do not remove user-defined metadata from documents. After downloading and storing files on disk, addressing data for high-level continuation relationships can be stored in two locations: 1. File system extended attributes, and 2. Extensible attributes in the file format.

[0384] For example, Table 4 below lists the addressing data storage method for changing the advanced continuation relationship in response to the change of the document storage location. It should be noted that the "cognitive element collection node" involved in Table 4 can be replaced by "advanced continuation relationship", and accordingly, the "addressing data of the cognitive element collection node" in Table 4 can be replaced by "addressing data of the advanced continuation relationship", and the same solution is applicable. Table 4

[0385] Referring to Table 4, when the document position changes, the storage method of the addressing data of the high-level continuation relationship can be updated based on the position change and the file displacement change.

[0386] In some embodiments, when there is a risk of a document being moved from an intranet to a non-intranet, the embodiments of the present application can monitor document access events and identify some possible storage media such as "save as new file" to a USB device, "file copied" to a USB device, "file moved" to a USB device, etc.

[0387] In some embodiments, if the document is of a confidential level or other security level, when a "network upload transmission" action is identified, which may be a risk of using the document on the predefined corporate intranet, the file is redirected to encrypt the metadata. <encryptedtag>The document that stores the addressing data of the updated high-level continuation relationship is given to the application for access. In this way, the document sent to the external network is in an encrypted and encapsulated state. Even unauthorized people cannot obtain the document content, or it is difficult to separate the document from the metadata such as the addressing data of the high-level continuation relationship due to the inability to decrypt it, and it is difficult to tamper with the metadata. If the document level is non-confidential, you can choose not to update it based on the decision. <encryptedtag>Encryption method, only <extensivetag>Way.

[0388] Step 120: Based on the cognize, update the high-level continuation relationship and / or the relationship between high-level continuation relationships corresponding to the cognize.

[0389] In the embodiment of the present application, the categories of cognitive relationship information predefined for a type of high-level continuation relationship are limited. If the cognitive relationship information of a cognizer does not belong to the category of cognitive relationship information constituting a high-level continuation relationship, the high-level continuation relationship cannot be newly created or modified.

[0390] In an optional embodiment, a high-level continuation relationship is created by a predefined high-level continuation relationship creation member cognate and a high-level continuation relationship continuation member cognate. Accordingly, in step 120, updating the high-level continuation relationship corresponding to the cognate includes creating the high-level continuation relationship and / or continuing the high-level continuation relationship.

[0391] In the embodiment of the present application, rules for determining high-level continuation relationships and / or relationships between high-level continuation relationships by cognitive elements may be predefined.

[0392] An optional implementation utilizes individual identifiers stored in extended attributes during the addressing data step to determine the document and its corresponding higher-level continuation relationship. For example, a document has a persistent FileID. A corresponding higher-level continuation relationship, EntityMirror.ActionBusiness, is established, with a one-to-one relationship between FileID and EntityMirror.ActionBusiness. The document entity flow relationship, FileFlowNode.ActionBusiness, can store the FileID or associate it with its corresponding EntityMirror.

[0393] In an optional implementation, the advanced continuation relationship includes document mirror awareness EntityMirror.ActionBusiness.

[0394] In an optional implementation manner, the high-level continuation relationship includes a content flow relationship ContentFlowNode.ActionBusiness (sometimes referred to as ContentFlowNode for short) and a document entity flow relationship FileFlowNode.ActionBusiness (sometimes referred to as FileFlowNode for short).

[0395] It is worth noting that this method of using only the original document individual attributes to determine the addressing data of the document and the corresponding high-level continuation relationship is similar in effect to using the individual identifier FileID stored in the extended attribute to uniquely mark a document. For simplicity, FileID is used as an example in this application.

[0396] It is worth noting that the entity relationship of various document individual attributes can be used to connect the combination of multiple document original attributes to represent the document mirror cognition EntityMirror.ActionBusiness, and this application does not limit this.

[0397] For ease of understanding, Table 4A below lists an example of a representation method of the document mirror cognition EntityMirror.ActionBusiness. The table uses the intra-entity relationship of each document's individual attributes to connect the combination of multiple document's original attributes to describe.

[0398] Table 4A

[0399] For example, a document entity can be described based on a combination of original document attributes such as document hash and document absolute path. For example, based on the location of the document, different categories of original document individual attributes are selected to establish a relationship with the high-level continuation relationship addressing data. For example, when the document is stored on a PC disk, the absolute path (document device path + file name) is used as the FileID, and a relationship between the FileID and the high-level continuation relationship document mirror cognition EntityMirror.ActionBusiness is established. When the document is transmitted over the network and leaves the PC, the document hash is used as the FileID, and a relationship between the FileID and the document mirror cognition EntityMirror.ActionBusiness is established. Since the document does not have original permanent attributes that can express the entity, the original attributes of the document change as the user operates on the document. For example, each time a document is edited, its hash may change. One method is to establish tuples corresponding to the document attributes in the database based on the intra-entity relationships of the document's individual attributes and the original attributes of the document, and to establish relationships between the tuples based on the intra-entity relationships of various document individual attributes. This is similar to recording various changes in document attributes within a document entity in the database, and recording changes in document attributes within the entity, such as recording multiple document hash versions of a document. This is essentially a document mirror cognition EntityMirror.ActionBusiness. This method is similar in effect to using the individual identifier FileID stored in the extended attribute and using EntityMirror.ActionBusiness to represent it in the database. For simplicity, this manual uses the individual identifier FileID stored in the extended attribute to correspond to EntityMirror.ActionBusiness.

[0400] In the embodiment of the present application, the document of the cognitive meta-operation has a unique corresponding document mirror cognition.

[0401] Due to user operations, individual document attributes often change. The following describes a method for establishing or maintaining individual document attributes and their corresponding high-level continuation relationships based on cognitive elements and cognitive relationship information.

[0402] In the method provided by the present application, a high-level continuation relationship and its corresponding creation member cognition element and continuation member cognition element are first predefined.

[0403] In some embodiments, in step 120 , the high-level continuation relationship corresponding to the cognize element and / or the relationship between high-level continuation relationships may be updated according to the type of the cognize element.

[0404] The type of a cognite can be determined based on the cognitive relationship information and / or the creation information of the higher-level continuation relationship corresponding to the cognite. In other words, when determining the type of a cognite, such as whether it is a member-creating cognite or a member-continuing cognite, the cognitive relationship information category and the creation information of the higher-level continuation relationship corresponding to the cognite can be considered. The creation information of the higher-level continuation relationship corresponding to the cognite can indicate that the higher-level continuation relationship corresponding to the cognite has been created, or that the higher-level continuation relationship corresponding to the cognite has not been created.

[0405] In one implementation, when a cognate's type matches a member cognate that is a creator of a higher-level continuation relationship, a new higher-level continuation relationship may be created for the cognate. Specifically, after obtaining a cognate, if the cognate's cognitive relationship information is a member cognate that is a creator of a higher-level continuation relationship corresponding to a source document, a new higher-level continuation relationship may be established, including a relationship between the new higher-level continuation relationship and an existing higher-level continuation relationship.

[0406] In another implementation, if the type of a cognate matches a continuation member cognate of a higher-level continuation relationship, the higher-level continuation relationship corresponding to the cognate can be continued. In other words, after obtaining a cognate, if the cognate's cognitive relationship information belongs to a continuation member cognate of a higher-level continuation relationship corresponding to the source document, the relationship between the higher-level continuation relationship and the document attribute is maintained.

[0407] In some embodiments, the high-level continuation relationship may include at least one of the following relationships: document mirroring cognition, document entity flow relationship, document content flow relationship, and folder cognition.

[0408] As an example, step 120 may specifically include: based on the folder attribute of the document of the cognitive meta-operation, updating the folder cognition corresponding to the folder attribute, wherein the folder cognition is a type of high-level continuation relationship.

[0409] As an example, step 120 may specifically include: associating the document operated by the cognitive element and the copy generated by the cognitive element with the same document entity flow relationship, wherein the document entity flow relationship is a high-level continuation relationship, and the cognitive element is a cognitive element that expresses the document mirror entity derivation relationship or the cognitive element that expresses the document being transmitted by the network.

[0410] In some embodiments, step 120 may specifically include: modifying the high-level continuation relationship corresponding to the cognitive element and the relationship between the high-level continuation relationships based on the cognitive relationship information and the non-document individual attributes.

[0411] In a possible implementation, after obtaining the epistem, the high-level continuation relationship #1 is updated based on the epistem relationship information, and the relationship between the high-level continuation relationship #1 and other high-level continuation relationships is updated based on the non-document individual attributes of the epistem.

[0412] As an example and not a limitation, the document entity flow relationship or document content flow relationship corresponding to the document of the cognitive meta-operation can be updated based on cognitive relationship information, and the relationship between the document entity flow relationship or document content flow relationship corresponding to the document of the cognitive meta-operation and the document entity flow relationship or document content flow relationship corresponding to the documents of other cognitive meta-operations can be updated based on non-document individual attributes.

[0413] In a possible implementation, after obtaining the cognition element, the high-level continuation relationship #2 corresponding to the non-document individual attribute of the cognition element is updated, and the relationship between the high-level continuation relationship #2 and other high-level continuation relationships is updated based on the cognition relationship information.

[0414] As an example but not limitation, the folder cognition corresponding to the folder attribute can be updated based on the folder attribute of the document of the cognitive meta-operation, and the relationship between the folder cognition corresponding to the folder attribute and the folder cognition corresponding to other folder attributes can be updated based on the cognitive relationship information, where non-document individual attributes include folder attributes.

[0415] It can be understood that the above-mentioned advanced continuation relationship #1 and advanced continuation relationship #2 are examples of advanced continuation relationships corresponding to cognitive elements.

[0416] For ease of understanding, Table 4B shows an example of a high-level continuation relationship, and its predefined create-member cognates and continuation-member cognates.

[0417] Table 4B

[0418] It should be noted that the advanced continuation relationship: Folder cognition Folder.ActionBusiness can include another implementation method: the cognition element establishes a one-to-one mapping relationship between PC documents and tuples in the database, and associates the database tuple with the cognitive attributes of the folder. Through mapping, the user's operation on any document in the folder, the generated cognition element can still be passed to the folder cognition, including the addition, modification, and deletion of documents. The effect is consistent with establishing the advanced continuation relationship Folder cognition Folder.ActionBusiness in the database.

[0419] In one embodiment, the addressing data of the high-level continuation relationship corresponding to the document attribute is directly stored in the extended attribute of the document, for example, EntityMirror.ID and FileFlowNode.ID are stored in the extended attribute of the document.

[0420] In another embodiment, in the step of establishing the document attributes and their corresponding higher-level continuation relationships, the individual identifiers stored in the document's extended attributes may be used. That is, a correspondence is established between the original document individual attributes of the document, the individual identifiers stored in the document's extended attributes, and the higher-level continuation relationships.

[0421] As an example, in response to the entity relationship EntityMirrorChanged for a document's individual attributes, a relationship is established between the original document's individual attributes and the individual identifier stored in the document's extended attributes, and a relationship is established between the individual identifier stored in the document's extended attributes and the corresponding high-level continuation relationship of the document attribute. Establishing the relationship between the original document's individual attributes and the individual identifier stored in the document's extended attributes includes creating a new FileID and storing it in the document's extended attributes. Establishing the corresponding high-level continuation relationship between the individual identifier stored in the document's extended attributes and the document attribute includes storing the FileID in the high-level continuation relationship, or determining a corresponding tuple FileMirror in a predefined location DB based on the FileID, and associating the high-level continuation relationship with the FileMirror, such as associating the EntityMirror.ID or FileFlowNode.ID.

[0422] As an example, in response to the intra-entity relationship EntMirrorRemain of the document's individual attributes, in response to the change of the original document's individual attributes, the correspondence between the changed original document's individual attributes and the corresponding high-level continuation relationship before the change is maintained unchanged. An optional method is, in response to the intra-entity relationship EntMirrorRemain of the document's individual attributes, in response to the change of the original document's individual attributes, the individual identifier stored in the corresponding document's extended attributes is maintained unchanged, and the relationship between the corresponding high-level continuation relationship and the individual identifier stored in the document's extended attributes is maintained unchanged. An optional method is, in response to the intra-entity relationship EntMirrorRemain of the document's individual attributes, in response to the update of the document's original attributes, the FileID is updated and stored in the document's extended attributes (i.e., the relationship between the updated original document attributes and the updated FileID is maintained), and the corresponding EntityMirror.ID is updated based on the updated FilID. Although the FIleID value and the EntityMirror.ID value have changed, the document still maintains the correspondence with the corresponding high-level continuation relationship EntityMirror.ActionBusiness.

[0423] In summary, in some embodiments, in step 120 , at least the individual identifier stored in the document extended attribute is used, and the individual identifier is used to determine the correspondence between the document and the high-level continuation relationship.

[0424] In some embodiments, the individual identifier includes at least one of the following: a document identifier, a document image identifier, and an identifier that enables a one-to-one mapping of a document to a document image.

[0425] In some embodiments, step 120 may specifically include: obtaining epistemic elements, and selectively updating high-level continuation relationships and / or relationships between high-level continuation relationships.

[0426] Selectively updating the high-level continuation relation, including selectively updating the cognitive attributes of the high-level continuation relation, the cognitive attributes including level categories;

[0427] Selectively updating relationships between superior continuation relationships, including updating the relationship degree and / or relationship direction of the superior continuation relationships.

[0428] That is, the relationship between high-level continuation relationships may include the degree of the relationship between high-level continuation relationships and the direction of the relationship between high-level continuation relationships.

[0429] In some embodiments, the direction of the relationship between high-level continuation relationships is determined based on the source and destination of the cognitive relationship information.

[0430] The attributes of a high-level continuation relationship can include original attributes and cognitive attributes. The original attributes are simply stored as the cognitive meta-attributes that make up the high-level cognitive relationship. For example, the attributes of an EntityMirror store the size, device ID number, folder attributes, application attributes, user attributes, device attributes, time attributes, etc. of the corresponding document. The cognitive attributes of a high-level continuation relationship are determined by multiple types of cognitive meta-attributes based on rules. They are important data for accurately determining the level and category of the high-level continuation relationship. Cognitive attributes include importance, etc.

[0431] As described in the table above, the cognitive attributes of FileFlowNode.ActionBusiness can continuously accumulate throughout the entire lifecycle of document storage, use, and network circulation. This reflects the continuous accumulation of multidimensional cognition carried by multiple cognitive elements acquired by the system's monitoring program within the same FileFlowNode.ActionBusiness after a document begins to flow, representing the cognitive information carried by the document flow.

[0432] As described in the above table, the cognitive attributes of the advanced continuation relationship of EntityMirror.ActionBusiness can be understood as data generated by the continuous accumulation of multiple operations (multiple cognitive elements) on a single document entity.

[0433] It should be noted that the process from obtaining the cognitive attributes of a simple high-level continuation relationship to obtaining the precise level category of the document corresponding to the high-level continuation relationship can be understood as a relationship where quantitative change leads to qualitative change. The level category of the high-level continuation relationship includes business attributes, which are a type of cognitive attributes of the high-level continuation relationship. Changes in the cognitive attributes of the high-level continuation relationship will only update the level category of the document corresponding to the high-level continuation relationship if certain conditions are met. For example, only when the attribute feature combination of FileFlowNode meets the predefined requirements will the level category of the document corresponding to FileFlowNode be changed. As an example, a cognitive element is obtained, and based on the addressing data of the high-level continuation relationship corresponding to the document, one or more corresponding high-level continuation relationships are updated, including selectively updating the cognitive attributes of the high-level continuation relationship; for example, if the document of the FileDerived category cognitive element is associated with EntityMirror.ActionBusiness and FileFlowNode.ActionBusiness, the attributes of both will be updated at the same time.

[0434] In one example, a cognitive element is obtained, and based on the addressing data of the high-level continuation relationship corresponding to the document, the relationship between the corresponding high-level continuation relationship and the existing high-level continuation relationship is updated; the updating of the relationship between the high-level continuation relationships includes updating the relationship degree and / or relationship direction of the high-level continuation relationship.

[0435] In some embodiments, selectively updating the relationship between the higher-level continuation relationships includes updating the relationship degree and / or relationship direction of the higher-level continuation relationships.

[0436] For example, different categories of basic events determine cognitive relationships. Based on rules, different levels of high-level continuation relationships can be determined. For details, see Table 10: Cognitive Elements, Updated High-Level Continuation Relationship Examples.

[0437] In one example, when the level category of the advanced continuation relationship #1 is changed, the business level category of the advanced continuation relationship #2 is selectively updated based on the relationship degree and / or relationship direction between the advanced continuation relationships.

[0438] In one example, a cognitive element is obtained and one or more high-level continuation relation cognitive attributes are selectively updated.

[0439] Another example, obtain a cognitive element and selectively update the relationship between high-level continuation relationships.

[0440] The selectivity includes determining one or more target high-level continuation relationship attributes to be updated based on the cognito element's cognito element set addressing data, and determining the high-level continuation relationship to be updated based on the cognito element's cognito element set addressing data and cognito attribute information. For example, a FileDerived class cognito element, based on its operation document, is associated with addressing data of multiple types of high-level continuation relationships (e.g., associated with EntityMirror.ActionBusiness and FIleFlowNode.ActionBusiness), so EntityMirror.ActionBusiness and FIleFlowNode.ActionBusiness may be updated simultaneously.

[0441] Table 5 shows the combination of cognitive relationships of multiple categories of basic events, which realizes the accumulation of cognitive attributes determined by basic events in the entire cycle of document mirroring.

[0442] Table 5

[0443] The following introduces an advanced continuation relationship: document mirror EntityMirror, and another advanced continuation relationship is FileFlowNode.

[0444] The document mirror EntityMirror.ActionBusiness can be understood as the accumulation of multiple operations on a single document entity. The document mirror EntityMirror can be called a mirror node or EntityMirror. The EntityMirror can store the document path FilePath.

[0445] For example, there are the FileDerive cognitive element for the document mirror entity derivation relationship, the DataInMotion cognitive element for the document network transmission relationship, and the EntityMirrorRemain cognitive element for the intra-entity relationship of the document's individual attributes. These cognitive elements contain addresses to the same predefined location or set of attribute combinations (such as the FileFlowNode stored in a database). These cognitive elements are based on decision-determined sets and can be understood as the continuous accumulation of multi-dimensional cognitive attributes (cognitive attributes determined by basic events) carried by the cognitive elements in the same FileFlowNode.ActionBusiness. This can reflect the document-based categories and business types in the user's mind, and embody category cognition through multi-user collaboration. In other words, the distribution of a document and all its copies generated by collaborative documents is based on the cognition embodied in predefined combinations of subject attributes.

[0446] In one example, the server aggregates document mirrors (EntityMirrors) using the FileID stored in the document's extended attributes. Based on the EntityMirrors' associations, the server then infers the collaborative relationships between different document entities. Different categories of cognitive attributes are accumulated within the same high-level continuation relationship (EntityMirror).

[0447] For example, the EntityMirror.ActionBusiness portion of the derived destination document is related to the EntityMirror.ActionBusiness of the derived source document, as shown in Table 6.

[0448] For each cognitive attribute determined by a basic event of different categories, the corresponding rule is selected to determine MultiRelationNode.ActionBussiness. For example, for the FilePath of the cognitive element, the corresponding sub-attribute of MultiRelationNode.ActionBussiness is selectively determined based on the FilePath rule. For the Time of the cognitive element, the corresponding sub-attribute of MultiRelationNode.ActionBussiness is selectively determined based on the Time rule.

[0449] Table 6 shows the acquisition of cognitive elements, the source document mirror cognitive EntityMirror.ActionBusiness based on the cognitive elements, and the selective determination of the destination document mirror cognitive EntityMirror.ActionBusiness based on the decision.

[0450] Table 6

[0451] The following describes the combination of multiple cognitive elements and the determination of FileFlowNode.ActionBusiness as an example in conjunction with Table 7.

[0452] Table 7: FileFlowNode.ActionBusiness property information and corresponding property descriptions

[0453] Step 130: When the updated high-level continuation relationship and / or the relationship between high-level continuation relationships meets the predefined decision, the level category of the document corresponding to the updated high-level continuation relationship is updated, wherein the predefined decision includes high-level continuation relationship features and their corresponding level categories.

[0454] In some embodiments, high-level continuation relationship features and their corresponding level categories may be based on manual settings.

[0455] As an example, the PC-side folders can be visualized on the server, and the folders can be divided into different levels and categories from the server.

[0456] For example, the first step: obtain the cognition element, obtain the high-level continuation relationship folder cognition: Folder.ActionBusiness, and display the document name and document properties in the folder on the PC in the server interface. Since the folder name can generally display its business and category, the document name can also display the business and category of the document. The second step: obtain user input (business, category) to modify the level category of the high-level continuation relationship Folder.ActionBusiness. The third step: based on the level category of the folder cognition, the client determines the extended attributes of the document associated with the folder cognition. For example, in the folder, a new document created by the user will be automatically classified into the level category of the folder cognition. It is worth noting that changing the level category needs to follow certain principles, such as the BLP rule. If the user marks a folder as public, when a confidential-level folder is stored in the folder, it should still maintain the confidential level.

[0457] An example of a high-level continuation relationship may have several boundary attribute features, and the boundary attributes are determined by a plurality of predefined subject attributes, and the subject attributes include devices, applications, and users. For example, a department boundary is composed of multiple predefined devices, and a FileFlowNode with member documents throughout the company has several department boundary attributes. An example is to determine multiple attribute features of a high-level continuation relationship, and the attribute features include but are not limited to category combinations of cognitive relationship information, folder attributes, subject attribute sequences, and application combinations, and predefine the attribute features of the high-level continuation relationship or the level categories corresponding to the attribute feature combinations. For example, the boundary of Department A is predefined to be composed of multiple devices, and documents flow through Department A, then FileFlowNode.ActionBusiness has the boundary attribute Department A. The department boundary attribute of FileFlowNode can reflect the flow of documents within the department, which can be called a department workflow.

[0458] As an example, the attribute characteristics of a high-level continuation relationship may include importance: auxiliary (ShareLevel=1), or primary (ShareLevel=2 or ShareLevel=3). For example, a document that is downloaded within a department boundary and is only read or edited, and is not analyzed and collaborated with others, can be understood as having an auxiliary importance within the department boundary, and its purpose is to serve documents with a "primary" importance. Combining the importance characteristics and the set boundary characteristics, the purpose of the document within the boundary can be determined. Only documents that are collaborated by multiple people can be understood as having a "primary" importance, with ShareLevel=2 for collaboration within the department and ShareLevel=3 for collaboration across departments. This is described in Table 7A below.

[0459] Table 7A

[0460] One idea of ​​this application is that there may be 1 million documents operated by an employee of Department A, and there may be only 50,000 FileFlowNode.ActionBusiness corresponding to the 1 million documents, that is, the boundary features of 50,000 FileFlowNode.ActionBusiness include Department A. The level category corresponding to the attribute feature is obtained by combining several attribute features. For example, only 200 attribute feature combinations may be needed to classify 50,000 FileFlowNode.ActionBusiness. A department always corresponds to a certain business, and the workflow feature combination may correspond to one or more types of business attributes (level categories). For example, Business A corresponds to 10 types of workflow features. By combining different attribute features of high-level continuation relationships and gradually screening to find out the corresponding one or more business relationships, the attribute features of the high-level continuation relationship can be classified into the corresponding one or more businesses. The following example steps are used to manually judge the business of the department's workflow features.

[0461] Feature acquisition 1: Obtain the boundary feature of the high-level continuation relationship FileFlowNode.ActionBusiness (including obtaining the corresponding department boundary feature consisting of multiple predefined devices);

[0462] Feature acquisition 2: Obtain the importance of the high-level continuation relationship FileFlowNode.ActionBusiness, for example, remove the FileFlowNodes with auxiliary importance and keep the FileFlowNodes with main importance.

[0463] Feature Acquisition 3: Obtain application combination features for the high-level continuation relationship FileFlowNode.ActionBusiness. For example, categorizing applications into two types, local PC applications and collaborative business applications, can filter out meaningless local applications, such as explorer.exe. The core purpose is to classify local applications into local applications with distinct business characteristics (such as VisualStudio.exe and Autocad.exe), local applications with less distinct business characteristics (such as Word.exe), collaborative applications with distinct business characteristics (such as OA, CRM, and ERP), and collaborative applications with less distinct business characteristics (such as Wechat.exe). Therefore, workflows within tens of thousands of departments can be divided into dozens or even hundreds of application combination features. For example, Application Combination Feature 1 (Word creation, OA collaboration) corresponds to 3,434 FileFlowNode.ActionBusinesses, Application Combination Feature 2 (Word creation, ERP collaboration) corresponds to 587 FileFlowNode.ActionBusinesses, and Application Combination Feature 3 (Word creation, CRM collaboration) corresponds to 1,547 FileFlowNode.ActionBusinesses.

[0464] Feature Acquisition 4: Obtain the subject attribute sequence features of the high-level continuation relationship FileFlowNode.ActionBusiness, for example, intra-departmental collaboration and cross-departmental collaboration. The boundary feature is Department A; the subject attribute sequence feature is Intra-departmental Collaboration; and the application combination feature is Word Creation and ERP Collaboration. This combination of three attributes corresponds to 587 FileFlowNode.ActionBusiness instances, indicating collaboration only among Leo, Philip, and Jack.

[0465] If the combination of the above four attribute features of FileFlowNode.ActionBusiness meets the first decision subset in the predefined decision, the first decision subset will classify thousands of documents corresponding to the 587 FileFlowNode.ActionBusiness as: logistics business.

[0466] If we further obtain the folder attributes of the corresponding users mentioned above, for example, the combination of the above four features corresponds to 587 FileFlowNode.ActionBusiness, which are distributed in the 10 folders of the three users Leo, Philip, and Jack. Since the folder name can generally reflect its business and category, for example, the corresponding folder names of Leo and Jack are: LeoPC\c:\Logistics Quote, LeoPC\c:\Logistics Supplier, JackPC:\D:\Logistics, indicating that the business corresponding to the application combination feature 2 is a logistics-related business. This can verify that the combination of attribute features based on the advanced continuation relationship meets the first decision subset in the predefined decision, and the correctness of the method for updating the level category of the document corresponding to the updated advanced continuation relationship. It is worth noting that the combination of multiple attribute features of the advanced continuation relationship, especially the combination with the corresponding folder attribute features, can help define the decision, such as what specific level category the combination of certain attribute features of the first decision subset should be classified into. Because the folder name can generally reflect its business and category, the document name can also reflect the business and category of the document.

[0467] It is worth noting that, for example, the attribute feature combination that the first high-level continuation relationship conforms to corresponds to multiple businesses. An optional method is to comprehensively consider the relationship between the first high-level continuation relationship and the second high-level continuation relationship, the attribute feature combination that the first high-level continuation relationship conforms to, and the level category combination of the second high-level continuation relationship, in accordance with the third decision subset in the predefined decision, and update the level category of the document corresponding to the first high-level continuation relationship.

[0468] For example, a new high-level continuation relationship has attribute feature combinations corresponding to three category levels (for example, human resources, finance, and administration). By leveraging the relationship between high-level continuation relationships (the importance of the primary high-level continuation relationship), the corresponding auxiliary FileFlowNode category is financial documents. Therefore, based on the predefined decision, the FileFlowNode is classified as a financial document. This is because when editing a primary document, users always reference documents related to the category. For example, when editing a financial document, there is a high probability that documents from other financial categories will be referenced.

[0469] Since this application is a positive feedback system, with the initial delineation of the level categories corresponding to the attribute features, the precise setting of the folder features, and the manual marking of the correction of the document level categories, the more accurate the level category division of the auxiliary FileFlowNode is, the higher its accuracy will be based on the correlation between them, as the relationship between the high-level continuation relationships is established.

[0470] Therefore, this method can greatly improve the recognition efficiency of document business. For example, after installing this system, it will run silently for one month. The 200 users in the department operated 100,000 documents, corresponding to 8,000 FileFlowNode.ActionBusiness, and 120 predefined attribute feature combinations including department boundary features (or called department workflow features). An administrator or someone familiar with the department's business can classify the 120 department workflow features into corresponding level categories, and then classify the 8,000 FileFlowNode.ActionBusiness corresponding to the 120 department workflows, that is, the number of documents that have completed business recognition reaches as many as 100,000 documents, and it only takes a short time. An optional method is to adjust and observe at regular intervals. After a period of time, you can understand all the business of the department and its corresponding document workflow features. In this way, the level category corresponding to the document can be accurately identified.

[0471] In one example, a cognitive element is obtained, one or more high-level continuation relationships are selectively updated, and if the high-level continuation relationship meets the attribute characteristics of a certain category of high-level continuation relationships, the high-level continuation relationship and / or the document corresponding to the high-level continuation relationship are classified into corresponding level categories.

[0472] For example, if the sequence of cognitive relationships of the basic events meets the predefined criteria, the high-level continuation relationship level category MultiRelationNode.ActionBusinessType is determined. In one embodiment of the present application, if the sequence of the subject attributes of a FileFlowNode meets the predefined criteria, a high-level continuation relationship level category is determined; and if the combination of multiple cognitive elements of a FileFlowNode meets the predefined criteria, a high-level continuation relationship level category is determined.

[0473] Changes to the level category of a high-level continuation relationship can update the level category of the documents associated with the high-level continuation relationship. This can be selectively updated based on rules. For example, if the level category is set to the same within a FileFlowNode, then when the high-level continuation relationship level category changes, the level category of all documents corresponding to that FileFlowNode will be changed to the updated level category.

[0474] The following describes in detail a method for determining the level category of a high-level continuation relationship in conjunction with Table 8.

[0475] Table 8 Combination update level category of cognitive element ActionBusinessType

[0476] In conjunction with the above example, in some embodiments, prior to step 130, attribute features of the high-level continuation relationship may be determined. The attribute features include at least one of the following: boundary features, the importance of the high-level continuation relationship, folder attributes corresponding to the high-level continuation relationship, the order of subject attributes corresponding to the high-level continuation relationship, and the application combination of the high-level continuation relationship. Step 130 may specifically include: when the attribute features of the high-level continuation relationship, or a combination of the attribute features of the high-level continuation relationship, meet a first decision subset of predefined decisions, updating the level category of the document corresponding to the updated high-level continuation relationship. The first decision subset includes the attribute features of the high-level continuation relationship and its corresponding level category, and / or a combination of the attribute features of the high-level continuation relationship and its corresponding level category.

[0477] As an example, if the boundary features and subject attribute order of the advanced continuation relationship meet the first decision subset, the advanced continuation relationship that meets the attribute feature combination is determined to be a predefined level category, and\or, the document corresponding to the advanced continuation relationship that meets the attribute feature combination is determined to be a predefined level category, and the subject attribute order includes the time series of the subject attributes.

[0478] As an example, if the boundary characteristics and importance of the advanced continuation relationship meet the first decision subset, the advanced continuation relationship that meets the attribute feature combination is determined to be a predefined level category, and\or, the document corresponding to the advanced continuation relationship that meets the attribute feature combination is determined to be a predefined level category.

[0479] As another example, if the boundary features and application combination methods of the advanced continuation relationship meet the first decision subset, the advanced continuation relationship that meets the attribute feature combination is determined to be a predefined level category, and\or, the document corresponding to the advanced continuation relationship that meets the attribute feature combination is determined to be a predefined level category.

[0480] As another example, if the combination of the boundary features, application combination methods and importance of the advanced continuation relationship of the advanced continuation relationship meets the first decision subset, the advanced continuation relationship that meets the attribute feature combination is determined to be a predefined level category, and\or, the document corresponding to the advanced continuation relationship that meets the attribute feature combination is determined to be a predefined level category.

[0481] As another example, if the combination of the boundary characteristics of the advanced continuation relationship, the application combination method, and the folder attributes corresponding to the advanced continuation relationship meets the first decision subset, the advanced continuation relationship that meets the attribute feature combination is determined to be a predefined level category, and\or, the document corresponding to the advanced continuation relationship that meets the attribute feature combination is determined to be a predefined level category.

[0482] In some embodiments, when the folder attribute of the updated high-level continuation relationship meets the second decision subset in the predefined decision, the level category of the document corresponding to the updated high-level continuation relationship is updated, and the second decision subset includes the folder attribute of the high-level continuation relationship and its corresponding level category.

[0483] In some embodiments, when the relationship between the first high-level continuation relationship and the second high-level continuation relationship, the attribute characteristics of the first high-level continuation relationship, and the combination of the level categories of the second high-level continuation relationship meet the third decision subset in the predefined decision, the level category of the document corresponding to the first high-level continuation relationship is updated, the first high-level continuation relationship is the high-level continuation relationship corresponding to the document of the cognitive meta-operation, and the third decision subset includes the attribute characteristics of the first high-level continuation relationship, the relationship between the first high-level continuation relationship and the second high-level continuation relationship, the combination of the level categories of the second high-level continuation relationship and its corresponding level categories.

[0484] In conjunction with Table 9, a cognitive element may contain multiple types of basic event cognitive relationships, modify multiple types of advanced continuation relationships, and modify the relationships between advanced continuation relationships. The following example uses the multiple document operations of the document family FileFlowNode as an example. It should be noted that the "cognitive element collection node" in Table 9 can be replaced with "advanced continuation relationship". Table 9

[0485] The following describes the specific method of obtaining cognitive elements, selectively modifying high-level continuation relationships, and relationships between high-level continuation relationships based on decisions.

[0486] One implementation method is to obtain cognitive elements, and the high-level continuation relations and the relationship categories between high-level continuation relations can be determined through a directed graph, while updating the vertices and edges of G(V, E). An optional storage format (triplet) is: (Fi, Fj, w), for example, where Fi represents the source MultiRelationNode, tj represents the destination MultiRelationNode, and W represents the weight. It should be understood that in a directed graph, the nodes of the graph can be document mirrors EntityMirror, and the edges represent the relationship between document mirrors EntityMirror-EntityMirror. The weight of the edge = the weight corresponding to the predefined cognitive relationship information category. In one example, the calculation model is a sparse matrix. Based on the <document-EntityMirror> of the cognitive element, find the node EntityMirror of the document graph.

[0487] In another directed graph, the node of the graph may be FileFlowNode. Based on the cognitive element, the node FileFlowNode of the document graph is modified, and based on the cognitive relationship information type of the cognitive element, the edge (E) is modified.

[0488] In some embodiments, updating the relationship between high-level continuation relationships includes establishing a direction of the relationship between the high-level continuation relationship corresponding to the source document and the high-level continuation relationship corresponding to the destination document based on the source document and the destination document of the cognitive relationship information.

[0489] In conjunction with Table 10, the cognitive element not only updates the level category of the advanced continuation relationship, but also may update the relationship between the advanced continuation relationships of the document. The following example illustrates this:

[0490] Table 10: Examples of relationships between epistemic elements and updated high-level continuation relationships

[0491] In one possible implementation, the aforementioned method for obtaining document cognition can be applied to hierarchically classify the document. For example, in response to a change in the level category of a first high-level continuation relationship, if the level category determination method for the first high-level continuation relationship is a high-confidence classification method, the level category of a second document or a second high-level continuation relationship determined using a low-confidence classification method can be updated based on the relationship between high-level continuation relationships. High-confidence level category classification methods may include: a decision based on the order of subject attributes of the high-level continuation relationship, a combination of multiple category cognates stored for the high-level continuation relationship, manual labeling, or level categories obtained from third-party applications.

[0492] In the step of determining the level category of the second document or the second high-level continuation relationship, methods for determining the level category of the high-level continuation relationship are compared.

[0493] The following describes in detail a specific implementation method for selectively updating the level category of a MultiRelationNode based on the method for determining the level category of a document.

[0494] In one implementation, a change in the level category of a first document changes the level category of a second document, and the step considers the degree of relationship between the first document and the second document; the step considers the credibility of the method for determining the level categories of the first and second documents.

[0495] In one implementation method, the first document and the second document are associated with the same MultiRelationNode, for example, the same FileFlowNode. According to the change in the hierarchical classification result of a member document associated with the FileFlowNode, the credibility of the level category determination methods of different documents associated with the FileFlowNode is compared, and its credibility is greater than the credibility of the level category determination methods of other documents. It is also compared whether the credibility of its level category determination method is greater than a threshold value, and then the level category of the second document is selectively updated.

[0496] In one implementation, the first document and the second document are associated with different MultiRelationNodes, and the degree and direction of the relationship between the MultiRelationNode associated with the first document and the MultiRelationNode associated with the second document, as well as the credibility of the level category determination method of the first and second documents, and\or the credibility of the level category determination method of the first MultiRelationNode and the second MultiRelationNode are taken into consideration.

[0497] EntityMirror.IdentifyType = "ShareAPP", for example, credibility = 80 points.

[0498] EntityMirror.IdentifyType = "SubjectOrder", credibility = 70 points.

[0499] A method for changing the classification of a member (a document) in a family tree: When a document is uploaded to a predefined app, such as a collaborative application like ShareBox, the classification of the data transmission from the ShareBox app (the classification obtained from a third-party application) is obtained. Because EntityMirror.IdentifyType = "ShareAPP", the confidence level is high (confidence = 80 points). Therefore, the classification of the document determined by the less reliable classification method with a close relationship may be updated based on rules.

[0500] In another possible implementation, a change in the level category of a first document changes the level category of a second document, and the step is based on rules and considers the scope of the second document that may be affected. In the step of determining the scope of the second document or the second high-level continuation relationship that may be affected, the relationship between high-level continuation relationships is used, including the degree of the relationship and the direction of the relationship. For example, the level category change of a FileFlowNode with ShareLevel=2 can be defined to only affect the document level categories (adjacent nodes) of other FileFlowNodes within the target range with a FolderBusiness relationship degree of 1 (the shortest path is 1). For example, a change in the classification result of a FileFlowNode with a higher ShareLevel value will affect all the corresponding documents, and the path correlation degree with all its documents is 1 (for example, in the same folder, the shortest path is 1), but the ShareLevel is lower than the ShareLevel value of the documents in the family tree, and then evaluate whether to change the level category of the target range documents.

[0501] In another possible implementation, the step of determining the scope of potentially affected second documents or second-level continuation relationships includes only considering the degree of relationship with the first document, and not considering the subsequent impact of changes to the level classification of the second document or the second-level continuation relationship. For example, if a change to the level classification of a first document changes the level classification of a second document, the subsequent impact of the change to the second document's level classification is no longer evaluated if the second document's level classification changes, thereby avoiding any ongoing impact.

[0502] In another example, a change in the level category of a FileFlowNode with ShareLevel=2 only affects the level category (adjacent nodes) of FileFlowNode with ShareLevel=1 created by the same user within the weight range of the relationship with the FileFlowNode. During the research process, the embodiment of the present application found that the document family (ShareLevel=1) on the PC accounts for more than 90% of the entire document family, that is, a large number of documents are not collaborated after creation. A common situation is that in order to edit the content of a document A.doc, a user may need the assistance of multiple documents. For example, 8 documents may be downloaded from the external network and 2 documents may be downloaded from the internal network. Then, through the time attribute related to the content usage time, it may be found that "Document A.doc" is the document with the most InDegree in the folder; similarly, through the "content reference relationship CopyContent", it is found that the user copied the content from the clipboard of 5 files in the folder to "Document A.doc", that is, "Document A.doc" is also the document with the most InDegree in "content reference relationship CopyContent".

[0503] Therefore, the storage and use of these 10 documents (auxiliary documents) may be to support the editing of the primary document family "Document A.doc". Therefore, within the folder scope, we can find the most important documents and document families based on the relationship type, in-degree, out-degree, and weight. The supporting documents can be understood as auxiliary documents and auxiliary document families.

[0504] After the main FileFlowNode described in "Document A.doc" is collaborated by the user, due to the high accuracy of the hierarchical classification through SubjectOrder, for example, it is classified as "intellectual property", so through the relationship between FileFlowNode-FileFlowNode, it can also be roughly determined that the category of these 10 auxiliary document families that have not yet been divided into level categories is "intellectual property".

[0505] For example, FileFlowNode.ActionBusinessType with ID=9688 = "Technical Document" (ShareLevel=1), and the closely related FileFlowNode.ActionBusinessType with ID=3918 (the FileFlowNode with the shortest path being 1) is changed from "Technical Document" to "Human Resources Document" (ShareLevel=2). Since the ShareLevel is higher, the classification method of its level category is more credible, so FileFlowNode.ActionBusinessType with ID=9688 is also affected and changed to "Human Resources Document".

[0506] For example, user A drafts a document and collaborates with multiple users. A categorized document (Enterprise Internal Control System.doc, FileFlowNode.ID=65685, Category: "Management System") is uploaded to the enterprise OA for company-wide sharing. A month later, user B downloads the document from the OA and several other documents from the external network, copies some of its content into a new document (Financial Management.doc, FileFlowNode.ID=65995), and collaborates on this document with multiple users. During the collaboration process, the FileFlowNode is marked as "Finance Category." Because FileFlowNode.ID=65995 has a ShareLevel=2, and FileFlowNode.ID=65685 also has a ShareLevel=2, the category of FileFlowNode.ID=65685 remains "Management System." Based on the decision, only closely related, less important, and uncategorized file families are affected. FileFlowNode documents with a ShareLevel=1 associated with FileFlowNode.ID=65995 are updated to "Finance Category."

[0507] In another possible implementation, the above-described method for obtaining document knowledge can be applied to centrally control document access. For example, access to or use of a document can be controlled based on a combination of the following: the first high-level continuation relationship corresponding to the document, the relationship between the first high-level continuation relationship and the second high-level continuation relationship, and the level category of the second high-level continuation relationship.

[0508] A key point of this application is that, rather than directly identifying the content of a document, such as controlling document access based on the content identification results, it leverages high-level continuation relationships and relationships between high-level continuation relationships. For example, controlling access to multiple auxiliary documents relies on the level classification of the primary document family FileFlowNode that is most closely related to the auxiliary documents and has a high-confidence classification method.

[0509] That is, controlling access to documents depends on the combination of three things: 1) the FileFlowNode (auxiliary FileFlowNode) corresponding to the document; 2) the relationship between FileFlowNodes (find the FileFlowNode of the high-confidence classification method that has the closest relationship with the auxiliary FileFlowNode); 3) the level category of the main document family FileFlowNode determined by the high-confidence classification method.

[0510] Document access includes opening a file, writing to a file, deleting a file, changing file permissions, changing file attributes, opening an email message in a mail store, deleting an email in a mail store, retrieving a document from a document management system, storing a document in a document management system, or any other act of accessing a document or document repository.

[0511] In another possible implementation, the aforementioned method for acquiring document knowledge can be applied to generate a document family distribution map. For example, based on the relationships between high-level continuation relationships, multiple high-level continuation relationships of different categories are aggregated to generate an audit map representing the distribution of documents of different categories on enterprise PC desktops. For example, a document audit map consisting of auxiliary document FileFlowNodes and primary FileFlowNodes can be constructed. Assuming an enterprise has 100 PCs, the document audit map might show that these 100 PCs store a total of 30,000,000 documents, corresponding to 1,000,000 FileFlowNodes. Of these, 80,000 are FileFlowNodes for intra-departmental collaboration and 20,000 are FileFlowNodes for cross-departmental collaboration. Important documents within the enterprise can be considered to be those corresponding to the 100,000 primary document family FileFlowNodes that have been collaborated on by multiple people. The corresponding relationships between the 900,000 auxiliary FileFlowNodes that have never been collaborated on and the primary document family FileFlowNodes that have been collaborated on are then displayed. This demonstrates the distribution of documents on user devices, corresponding to these closely related FileFlowNodes at different levels.

[0512] For example, FIG3 is a schematic diagram of an audit drawing of a specific document.

[0513] Optionally, in some embodiments, the user's interests may be recorded, and appropriate documents may be automatically recommended to the user based on the results of document classification.

[0514] For example, a user's long-term interests (User.LongInterest.ActionBusiness) and current interests (User.NowInterest.ActionBusiness) are recorded based on time. For example, user Leo is currently working on three document families: FileFlowNode=156898, FileFlowNode=565665, and FileFlowNode=23645, all of which are "Technology" and "Project X." "Technology" and "Project X" are stored as Leo's current interests. When Leo opens an enterprise content management (ECM) document collaboration application, he or she may search for content related to "Project X." A feasible approach is to transfer the stored User.NowInterest.ActionBusiness to the ECM. This eliminates the need for the user to search for content related to "Project X." Instead, the ECM recommends content related to "Project X" to Leo. If no suitable content is found, the user can search again. This improves the user experience and increases the efficiency of ECM usage.

[0515] Optionally, in some embodiments, multiple high-level continuation relationships of user operations on documents and applications within a predefined time period can be obtained. Based on the boundary features of the high-level continuation relationships, such as the boundaries of different departments, one or more audit drawings can be generated. These audit drawings can be used to illustrate the main work distribution of department members or a specific user during working hours. For example, Leo worked on Project A for 4 hours.

[0516] Optionally, in some embodiments, high-level cognitive relationships determined by normal working behaviors are predefined and input into a predefined knowledge determination model or a data-driven model. Multiple cognitive elements can also be obtained to form multiple high-level continuation relationships, and the high-level relationships are input into a predefined knowledge determination model or a data-driven model to determine whether the behavior is normal.

[0517] Optionally, in some embodiments, if a document stored in a folder on a PC flows to the document management software ECM, and the location or category of the document stored in the ECM is obtained, the extended attributes of the document or folder can be updated to establish a correspondence between the PC folder and the ECM storage location, for example, it can be automatically synchronized to the corresponding location of the ECM.

[0518] Optionally, in some embodiments, a mapping relationship is established between individual document attributes and one or more documents in the document management software (ECM), such as a mapping relationship between collaborative documents. After a user edits document A locally, the edited content can be immediately synchronized to the corresponding document in ECM. This eliminates the need for users to collaborate on documents in ECM, allowing them to edit locally and automatically synchronize, resulting in a better user experience.

[0519] Optionally, in some embodiments, since a mapping is established between PC folders and ECM folders in the document management software, users can upload documents to the corresponding folder in ECM by right-clicking and clicking "Upload." This eliminates the need for users to drag and drop documents to the ECM network drive or to open a web browser, select the corresponding folder, and then click the "Upload" button to upload, thereby providing a better user experience.

[0520] The decision center can be a knowledge-driven model or a data-driven model, which is not specifically limited in the present embodiment. The knowledge-driven model and the data-driven model are described in detail below.

[0521] The knowledge-driven model is a relationship extraction model based on machine learning and deep learning methods. It can automatically learn knowledge from data that has been pre-labeled by experts and data labelers, thereby selectively controlling document operation events and realizing entity and relationship preservation.

[0522] Data-driven models are based on the experience of domain experts. They can design some rules or patterns and add them to the model to enable the model to quickly acquire knowledge. This can be implemented based on rules, patterns, and statistical methods.

[0523] The control rules used in the examples of this application include those based on the ABAC model. These control rules comply with the XACML standard. For example, when a document operation event occurs, the decision center can be queried to implement selective control of the application.

[0524] The following describes the evaluation process of the rules used in the control step of the application:

[0525] 1. Obtain the monitoring results of document operation events and collect information about events that require evaluation rules, including the event name, the value or attribute information of the subject and object corresponding to the event, and any information required by the evaluation rules, such as time information, network information, and other environmental information.

[0526] 2. Select an applicable policy. A rule query input consists of the identification information and attribute information of the event and its subject (user, behavior, application), the identification information and attribute information of the object (including source path, target path, extended attributes of the document being operated on, and other information), category information such as the environment, and related attributes to form a query request.

[0527] 3. Evaluate control rules that comply with the XACML standard. Based on the selected control sub-rules, obtain the other information required to complete the rule evaluation. If specific variables are used in the control rules, it is necessary to obtain the definition of the corresponding control rule variable interpretation. Substitute the entity identification information and entity attribute information according to the description of the definition of the control rule variable interpretation, thereby obtaining the information required to complete the rule evaluation using the control rule variable interpretation. Substitute the conditions expressed by attributes and grammar in the rule to determine whether the conditions are met. If the conditions are met, the rule is determined to be relevant. If a rule containing a conditional expression is matched, the result of the rule evaluation will include a result part of the rule containing the conditional expression. The result part of the rule often includes an action.

[0528] The actions in the decisions that meet the requirements are stored and used to implement application control. The actions may include, but are not limited to: blocking or changing the application's functions, after the user directly or indirectly calls but before the application's functions are executed, such as by controlling the kernel's input / output request packet (IRP), so that the user and / or the application interprets this situation as a specific operating system service failure or a hardware device application failure, and the data operation event will no longer be executed; or changing, replacing, removing, hiding, or obscuring one or more parts (or all) of the results to be presented to the user, changing, replacing, removing, hiding, disabling, or obscuring one or more operable objects or text fragments, such as transforming the addressing data storage method of a high-level continuation relationship, or performing a specified operation.

[0529] For illustrative purposes, only one rule is evaluated in this embodiment. In practice, rule evaluation can select one or more rules related to the behavior to determine whether the application behavior needs to continue executing or be controlled by additional actions. Furthermore, rule evaluation can include more than one rule. If the rule conditions are met, the rule's control result is executed. If not, the next control sub-rule is matched. It should be noted that multiple control sub-rules may meet the conditions. When the preconditions of more than one rule are met during rule evaluation, one or more combination algorithms must be used to combine the rule results of the evaluated rules to form the final rule result. An alternative implementation involves the control language being based on the XACML (eXtensible Access Control Markup Language) standard. XACML provides a rule conflict resolution method and utilizes conflict avoidance algorithms to ensure the certainty of the rule control system evaluation results. For example, when an application's access to a document is monitored, a rule query may find two rules matching the application's access behavior. After rule evaluation, if one rule's result is a deny and the other's result is an allow, then, according to the rule conflict resolution method, the result returned to the query is a deny, thus resolving the rule conflict. The final rule result is then returned to the rule query module. This final policy result usually contains an effect <effect>, optional 0 or more instructions <obligation>.

[0530] For example, the definition of control strategy number 1 is listed below:

[0531] Application conditions: Any explorer.exe program ( <application>image_path==[*explorer.exe] < / application> )

[0532] Document operation event category: "File created" ( <event> event_name==[create] < / event> )

[0533] Object condition: Any document name ( <resource> file.destination.path==[*] < / resource> )

[0534] If the above conditions are met,

[0535] Result 1 is: Allow the application to continue executing ( <result> allow < / result> );

[0536] Result 2 is: perform the specified operation <document original attribute-addressing data of advanced continuation relationship>

[0537] <obligation_resource> file.destination< / obligation_resource>

[0538] According to the rules, create FileFlowNode.ID:

[0539] <fileflownode>Create.FileFlowNode.ID< / FileFlowNode.ID>

[0540] According to the rule, create EntityMirror.ID: <entitymirror> Create.EntityMirror.ID < / entitymirror>

[0541] < / Addressing data of the original document attributes - advanced continuation relationship>

[0542] Result 3 is: Execute the specified operation: Store the specified metadata <Addressing data of the original document attributes - advanced continuation relationship>, that is, store the addressing data such as EntityMirror.ID and FileFlowNode.ID to the specified location in the document extension attributes.

[0543] <filetag>

[0544] <File.EntityMirror>File.destination.EntityMirror.ID< / File.EntityMirror>

[0545] <File.FileFlowNode>File.destination.FileFlowNode.ID< / File.FileFlowNode>

[0546] <obligation_resource>file.destination< / obligation_resource>

[0547] < / filetag>

[0548] For example, the following combinations of control strategies are listed:

[0549] Next, in conjunction with FIG. 4, taking the document operation event that the document is newly created as an example, combined with the above control strategies, a specific implementation method of the cognitive element is described in detail.

[0550] FIG. 4 is a schematic flowchart of a method for a "document newly created" cognitive element provided by an embodiment of the present application. As shown in FIG. 4, the method may include steps 410-460, and steps 410-460 are described in detail below.

[0551] Step 410: The customer performs an operation of creating a new document on the client.

[0552] As an example, a user designated as a manager performs a right-click new document operation on windows and creates a new document named "C:\X-P\New doc document.doc".

[0553] Step 420: The client outputs an event of "<Document newly created>".

[0554] As an example, the application monitoring program installed in the Windows system client outputs basic events, a combination of basic events, and outputs the "<document is newly created>" document operation event, including subject attributes (application name (here image_path == [explorer.exe])), event category (here event_name == [create]), object attributes (here file.source.path == [], file.destination.path == [C:\XP\New doc document.doc]), and uses the collected data as policy query information.

[0555] Step 430: The decision center receives the query information, performs a policy evaluation on the stored related policies, and outputs the policy effects.

[0556] As an example, the decision center receives the query information and performs a policy evaluation on the stored relevant policies. The policy evaluation includes collecting relevant data required for the policy evaluation. In this example, it is assumed that only the above policies are evaluated. In this example, the query matches the condition in the first policy ( <id> 1< / id> , then the policy consequences specified in the policy will be adopted. In this case, the policy evaluation produces a policy effect of ALLOW ( <result> allow < / result> ) and execute the specified Obligation. The decision center calls the Obligation handler to execute the Obligation task specified in the policy and returns the policy effect to the application monitoring event module.

[0557] Step 440: Determine whether the cognitive element "<document is newly created>" is included.

[0558] In this case, <document original attribute-addressing data of advanced continuation relationship> obligation module, <filetag>The modules are <obligation>Processor call. Among them, <document original attribute - addressing data of advanced continuation relationship> obligation creates a new EntityMirror.ID according to the rules. <filetag>, storing addressing data for high-level continuation relationships, such as EntityMirror.ID and FileFlowNode.ID, in the extended attributes of the destination document. Based on the location attributes of the destination document, the corresponding metadata storage method is adopted. For detailed metadata storage methods, please refer to the description above and will not be repeated here.

[0559] Step 450: The application monitor interceptor receives the policy consequence including the policy effect from the decision center. Result The application allows the operation to continue.

[0560] Step 460: Execute the application code that implements the file creation operation.

[0561] In this application example, explorer.exe executes the application code that implements the new file operation.

[0562] After receiving the cognitive element, including EntityMirror.ID, FileFlowNode.ID, cognitive attributes determined by the basic event, and cognitive relationship of EntityMirrorChanged, the server performs the following operations:

[0563] 1. In the database, based on the decision, create an EntityMirror (including determining EntityMirror.ActionBusiness) and a FileFlowNode (including FileFlowNode.ActionBusiness). The document-related relationships can be stored in the predefined relationship storage location associated with the EntityMirror.

[0564] 2. In the DB, update the relationship between high-level continuation relationships, such as updating the FileFlowNode-FileFlowNode relationship, such as updating the vertices and edges of G(V, E) at the same time.

[0565] 5, taking the "file is copied" cognitive element as an example, combined with the above control strategy, a specific implementation method for maintaining the correspondence between the original attributes of the document and the high-level continuation relations such as MultiRelationNode is described in detail.

[0566] Figure 5 is a schematic flow chart of a method for maintaining a correspondence between a document and a MultiRelationNode according to an embodiment of the present application. As shown in Figure 5, the method may include steps 510-540, which are described in detail below.

[0567] Step 510: The client executes a file copy operation on the client.

[0568] As an example, a user designated as a manager copies "C:\XP\Demo1.doc" to the destination document: "D:\Knoo\Demo1.doc" on Windows.

[0569] Step 520: The client outputs a "document copied" event.

[0570] As an example, the application monitoring program outputs basic events, a combination of basic events, and a "<document copied>" document operation event, and uses the collected data as policy query information.

[0571] Step 530: The decision center receives the query information, performs a policy evaluation on the stored related policies, and outputs the policy effects.

[0572] Similarly, the decision center receives the query information, selects the appropriate policy, and the policy evaluation produces the policy effect of ALLOW ( <result> allow < / result> ) and execute the specified Obligation. The decision center calls the Obligation handler to execute the Obligation task specified in the policy and returns the policy effect to the application monitoring event module.

[0573] Step 540: Store and maintain the correspondence between the document and the document data relationship.

[0574] As an example, in this case, <document original attribute-addressing data of high-level continuation relationship> obligation module, <filetag>The modules are <obligation>Processor call. Among them, <Document original attribute - addressing data of advanced continuation relationship> obligation creates a new EntityMirror.ID of the destination file (EntityMirror.ID=989665). <filetag>, associates the addressed data of the high-level continuation relationship with the destination file.

[0575] After receiving the cognitive element, including EntityMirror.ID, FileFlowNode.ID, and cognitive attributes determined by the basic event, the server performs the following operations:

[0576] 1) It includes deriving FileDerive=(656556:989665) based on the document mirror entity, and determining the destination document mirror EntityMirror.ActionBusiness, EntityMirror.ID=989665 and other properties based on the source document mirror EntityMirror.ActionBusiness.

[0577] 2) Based on the decision, determine <fileflownode-fileflownode>, update FileFlowNode.ActionBusiness,

[0578] 6 , taking the document operation event of a document being uploaded as an example, and combining the above control strategy, a specific implementation of a cognitive element containing a document operation event is described in detail.

[0579] Figure 6 is a schematic flow chart of a method for storing cognitive elements in a document uploaded according to an embodiment of the present application. As shown in Figure 5 , the method may include steps 610 - 650 , which are described in detail below.

[0580] Step 610: The client performs a document upload operation on the client.

[0581] As an example, a user designated as a manager uploads the D:\Rstt\Demo1.doc file (File.Source.EntityMirror.ID=989665) to ShareBox.exe on a Windows system.

[0582] Step 620: The client outputs the "<document uploaded and network information obtained>" event.

[0583] As an example, the API call caused by the user operation is intercepted by the application monitoring program installed in the Windows system client, and the basic event is output. After possible merging and combination, the "file is uploaded and network information is obtained" and event-related attributes are output, including the subject (application name, here is image_path == [ShareBox.exe])), event (here is event_name == [upload]), document attributes (here is file.source.path == [D:\Rstt\Demo1.doc], file.destination.path == []), and the collected data is forwarded to the decision center as policy query information.

[0584] When the "Select File" base event occurs, it's called by a network application, implementing document redirection. This involves copying the source file, D:\Rstt\Demo1.doc, to a temporary folder. The data needed to identify the MultiRelationNode, such as EntityMirror.ID=235656, is generated. The new document, which stores the addressing data for the higher-level continuation relationship, is then transferred to the network application.

[0585] Step 630: The decision center receives the query information, performs a policy evaluation on the stored related policies, and outputs the policy effects.

[0586] As an example, the decision center receives the query information and evaluates the application name, document attributes, event attributes and stored related policies submitted by the query. In this example, it is assumed that only the above policies are evaluated. In this example, the query matches the conditions of the 5th policy ( <id> 5< / id> , the policy consequence specified by the policy will be adopted. In this case, the policy evaluation produces a policy effect of ALLOW ( <result> allow < / result> ) and 2 Obligation( <obligation>The decision center calls the Obligation handler to execute the Obligation task specified in the policy and returns the policy effect to the application monitoring event module.

[0587] Step 640: Determine the metadata corresponding to the "document uploaded" event.

[0588] In this case, the <document original attribute-addressing data of the advanced continuation relationship> obligation module, <filetag>The modules are called <obligation>The processor is called. Because the corresponding redirection logic has been executed when the network application "selects a file", the new EntityMirror.ID is implemented in advance.

[0589] Step 650: The application monitor interceptor receives the policy consequence including the policy effect from the decision center, resulting in the application allowing the operation to continue.

[0590] Since the cognitive element containing the document operation event "a document is uploaded and network information is obtained" has multiple categories, the server performs the following operations:

[0591] Since it belongs to the DataInMotion category cognitive element, based on decision-making, cognitive attributes are determined based on the multi-dimensional basic events such as AppBusiness and UserBusiness carried by the cognitive element. Based on the decision center, a new EntityMirror is created in the DB, and FileFlowNode.ActionBusiness is selectively updated. <fileflownode-fileflownode>relation.

[0592] The cognitive attributes of the epistemic element and the high-level continuation relationship will only change the level category (ActionBusinessType) of the high-level continuation relationship when it is consistent with the decision. Based on the ActionBusinessType of Demo1.doc, the high-confidence classification method is changed to selectively change the level category of multiple documents associated with the same MultiRelationNode based on rules; and selectively change the level category of multiple documents associated with different MultiRelationNodes. The following uses the ShareAPP tag scenario as an example to describe how to selectively update the ActionBusinessType of the second document based on the decision center when the ActionBusinessType of the first document changes.

[0593] For example, when the Demo1.doc file (File.Source.EntityMirror.ID=235656) is uploaded to ShareBox.exe, the client has predefined communication rules with the third-party application ShareBox.exe and obtains the upload path and corresponding information. For example, it obtains "Demo1.doc Product Document\Product Design, ActionBusinessType: Product Document, Level: Internal, Master Document Family." This means that the server detects the change in the ActionBusinessType of the first document, EntityMirror.IdentifyType="ShareAPP."

[0594] In one case, the EntityMirror.IdentifyType of all documents in the <Document-FileFlowNode> associated with EntityMirror.ID=235656 is set to "Keyword Identification" with a credibility of 5 points. Since the IdentifyType of the changed document Demo1.doc is set to ShareAPP, which is higher than the current value, the .ActionBusinessType of all documents associated with this FileFlowNode is updated to "Product Document, Level: Internal, Master Document Family."

[0595] In one case, the degree and direction of the relationship between the high-level continuation relationships are predefined to determine the scope of the second document, based on <fileflownode-fileflownode>, update the member document FolderBusiness relationship degree of the FileFlowNode associated with the document to 1 (the shortest path is 1), and multiple FileFlowNode.ActionBusinessType with ShareLevel=1 and the ActionBusinessType of the documents associated with these FileFlowNodes, then some document levels are updated to "Product Document, Level: Internal, Auxiliary Document Family".

[0596] The following takes the manual marking scenario as an example to describe how to selectively update the ActionBusinessType of the second document based on the decision when the ActionBusinessType of the first document changes.

[0597] Similarly, for example, when a document in a FileFlowNode is manually marked as a "K project" by the user, if the other documents in the FileFlowNode to which the document belongs are classified by keyword recognition, since the credibility is lower than the credibility of the user's manual marking, the ActionBusinessType of the other documents in the FileFlowNode is changed to "K project". This method can improve efficiency. For example, if a document has 20 copies, it only needs to be manually marked once, and the other 19 derivative copies of the document associated with the same FileFlowNode can be automatically marked. Similar to the relationship between high-level continuation relationships, here is based on <fileflownode-fileflownode>Based on the relationship degree and relationship direction with the FileFlowNode, other predefined FileFlowNodes and their associated multiple documents are updated, and the ActionBusinessType of these documents is updated. Several unclassified documents whose relationship degree with the FileFlowNode is within the predefined range are also updated to "K project, auxiliary document".

[0598] 7 , taking FileFlowNode as an example, a specific implementation of how to determine a FileFlowNode with an associated relationship according to a decision center is described in detail.

[0599] Figure 7 is a schematic flow chart of a method for determining associated FileFlowNodes provided by an embodiment of the present application. As shown in Figure 7, the method may include steps 710-760, which are described in detail below.

[0600] Step 710: A new document is created, and the <document-document family (FileFlowNode)> is determined. <fileflownode-fileflownode>.

[0601] For example, a user designated as a manager creates a new document D:\Project DF\Craft1.doc file on a Windows system, with EntityMirror.ID=7565612 and FileFlowNode.ID=8651651, and edits the document.

[0602] Step 720: The document is downloaded and the <Document-FileFlowNode> is updated <fileflownode-fileflownode>.

[0603] As an example and for better editing of this document, two documents, namely External1.doc and External2.ppt, were downloaded from the external network, and one document, namely Internal1.doc, was downloaded from the internal network. All of them were stored in the D:\Project DF\ folder.

[0604] Step 730: The application monitoring program output a cognitive element determined as "file downloaded".

[0605] Step 740: The server obtained the cognitive element determined as "file downloaded", updated the cognitive element set node (such as the advanced continuation relationship), and updated the relationship between the cognitive element set nodes (such as the relationship between the advanced continuation relationships).

[0606] It is determined that the FileFlowNodes corresponding to the above three downloaded documents in <Document-FileFlowNode> are as follows:

[0607] <D:\Project DF\External1.doc ---- FileFlowNode.ID = <8686661>;

[0608] <D:\Project DF\External2.doc ---- FileFlowNode.ID = <1246565>;

[0609] <D:\Project DF\Internal1.doc ---- FileFlowNode.ID = 343434>;

[0610] The server obtained the cognitive element determined as "file downloaded", which belongs to the cognitive element of the DataInMotion category.

[0611] That is, the EntityMirrorChanged of the entity relationship of the document individual attributes was obtained, the FolderBusiness cognition was accumulated, and the FileFlowNode - FileFlowNode relationship was updated. Based on FolderBusiness, the FileFlowNode - FileFlowNode was determined. Since they are all stored in a folder D:\Project DF\, the shortest path of their document mirror images is 1, and the degree of their FolderBusiness relationship is 1.

[0612] Step 750: The application monitoring program output a cognitive element determined as "file edited".

[0613] For example, if a user edits the file D:\Project DF\Craft1.doc and also edits three documents in D:\Project DF\, the monitoring detects the "File Edited" cognitive element, FileFlowNode.ID=8651651. Furthermore, the "File Edited" cognitive elements, FileFlowNode.ID=8686661, FileFlowNode.ID=1246565, and FileFlowNode.ID=343434, are obtained.

[0614] Step 760: The server obtains the cognize element determined as "file is edited", updates the cognize element set nodes (eg, high-level continuation relationships), and updates the cognize element set node relationships (eg, relationships between high-level continuation relationships).

[0615] The server obtains the cognitive element that determines "the file is edited" and updates <Document-FileFlowNode> based on the time attribute contained in the cognitive element. <fileflownode-fileflownode>.

[0616] 8 , taking the four FileFlowNodes shown in FIG8 as an example, a specific implementation method of mutual influence between FileFlowNodes on the classification results of documents will be described in detail.

[0617] Figure 8 is a schematic flow chart of a method for classifying documents that influence each other between family trees according to an embodiment of the present application. As shown in Figure 8 , the method may include steps 810-850, which are described in detail below.

[0618] For example, Table 11 lists the FileFlowNode.ActionBusinessType predefined by the server.

[0619] Table 11 Predefined FileFlowNode.ActionBusinessType

[0620] Step 810: The first client performs a document uploading operation.

[0621] As an example, user Leo (first user) designated as a manager creates the file D:\Project DF\Craft1.doc on the Windows system, FileFlowNode.ID=8651651, and Leo (first customer) sends it to "Jack" (second user) through the wechat.exe user for collaboration.

[0622] Step 820: Based on the application monitoring program of the first client's personal computer (PC), a document operation event of "file uploaded and network information obtained" is output.

[0623] As an example, the monitoring program on the PC of user Leo outputs the cognitive element "the file is uploaded and network information is obtained";

[0624] Step 830: The server obtains the "file is uploaded and network information is obtained" cognition element and modifies the family tree (FileFlowNode).

[0625] Step 840: The application monitoring program of Jack's personal computer (PC) outputs the "file is downloaded and network information is obtained" cognition element.

[0626] Step 850: The server obtains the "file is downloaded and network information is obtained" cognition element and modifies the family tree (FileFlowNode).

[0627] Because the FileFlowNode collaborates with multiple users, the ShareLevel changes from 1 to 2. Since the order of the users of the FileFlowNode conforms to the predefined "FileFlowNode.UserOrder="Leo", "Jack"", the FileFlowNode.ActionBusinessType = "X Project".

[0628] Since FileFlowNode.ActionBusinessType = "X Project" has changed, based on <fileflownode-fileflownode>Since the degree of relationship with FileFlowNode.ID=8651651 is within a predefined range (for example, determined by multiple factors such as the shortest path of FolderBusiness being 1), the three FileFlowNodes (FileFlowNode.ID=8686661, FileFlowNode.ID=1246565, and FileFlowNode.ID=343434) all have ShareLevel=1. Therefore, according to the decision center, the three FileFlowNode.ActionBusiness can be updated to "X Project".

[0629] In another possible implementation, the above-described method for acquiring document knowledge can be applied to network recognition. For example, a combination of a document's classification and network data transmitted over the network can be determined, and based on the document's classification and the network data transmitted over the network, the document and / or its classification can be transmitted to a predefined network device.

[0630] Optionally, in some embodiments, due to encryption protocols such as https and private application protocols such as WeChat, it is difficult for current IDS, firewalls and other devices to obtain the document content transmitted in the IP data packet. Since the present application can accurately obtain the business type of the document, such as the FileFlowNode.ActionBusinessType associated with the document, as well as the URL, IP address, application name and other information where the document is uploaded, the present application can construct an "IP packet content" combination: (the document's ActionBusinessType, IP and other network data), and send the "IP packet content" to the network device.

[0631] One implementation method is to communicate with network devices such as firewalls, IDS, and IPS, and perform the following operations when a document is uploaded by a client to an internal or external server:

[0632] 1): The user designated as the manager uploads the D:\Rstt\Demo1.doc file (File.Source.EntityMirror.ID=989665) to wechat.exe on the Windows system.

[0633] 2): The API call caused by the user operation is intercepted by the application monitoring program installed in the Windows system client, which outputs the basic event. After possible merging and combination, it outputs "the file is uploaded and network information is obtained" and event-related attributes, including the subject (application name, here is image_path == [ShareBox.exe]), event (here is event_name == [upload]), document attributes (here is file.source.path == [D:\Rstt\Demo1.doc], file.destination.path == []), and forwards the collected data to the decision center as policy query information.

[0634] 3): The decision center receives the query information and evaluates the application name, document attributes, event attributes, and stored related policies submitted by the query. In this example, it is assumed that only the above policies are evaluated. In this example, the query matches the conditions of the 5th policy ( <id> 5< / id> , the policy consequence specified by the policy will be adopted. In this case, the policy evaluation produces a policy effect of ALLOW ( <result> allow< / result> ) and 3 Obligation( <obligation>The decision center calls the Obligation handler to execute the Obligation task specified in the policy and returns the policy effect to the application monitoring event module.

[0635] 4): In this case, the <document original attribute-addressing data of the advanced continuation relationship> obligation module, <filetag>obligation module, <ip-actionbusiness>The obligation modules are <obligation>Processor call. <ip-actionbusiness>Once invoked, data is transmitted to a network device at a predefined IP address. This data includes the destination network information for the "file uploaded and network information obtained" action: URL, IP address, application name, as well as the level and category of the uploaded source file. Upon receiving this transmitted data, network devices like IDSs no longer need to identify the document's level or category through keyword or content recognition, nor do they need to parse proprietary protocols like WeChat, enabling more refined network control.

[0636] To address the complex scientific challenge of business identification of unstructured data, this application has developed a relatively complete cognitive combination management system for documents stored on PCs, scientific discoveries for identifying documents through multiple cognitive combinations, basic theoretical methods, and engineering implementation systems. The application has three major scientific discoveries:

[0637] The first is the cognitive element generation system. In view of the heterogeneous, lack of structure, and volatile document attribute characteristics of unstructured data (hereinafter referred to as documents), a multi-dimensional cognitive aggregation method of "application basic event layer and cognitive element layer" is revealed. Document operation events are used to represent one or more types of cognitive relationship information. The user's operation on the document is abstracted as a cognitive element. Multiple cognitive elements are aggregated into high-level continuation relationships based on rules for accumulation and continuation of cognition. This cognitive element, based on the aggregation of addressing data of high-level continuation relationships corresponding to document attributes, realizes the aggregation of multi-dimensional cognition scattered across different PC devices to a predefined location server, and realizes the aggregation of cognition occurring at different times for the same document. For example, different categories of cognitive elements such as DataInMotion and DataInUse realize the orderly aggregation of multi-dimensional cognition occurring at different times for the same document.

[0638] The second is a system that selectively combines multiple categories of cognitive elements to determine the level categories (including business) of high-level continuity relationships. This system is composed of high-level continuity relationships constructed based on cognitive elements and the relationships between high-level continuity relationships. This cognitive accumulation system contains the flow characteristics of documents and the relationships between these flow characteristics.

[0639] Theoretically, multiple categories of high-level continuation relationships and the relationships between them can generate a complex and large graph. This application selects predefined categories of cognitive elements for selective combination to determine meaningful subgraphs such as FileFlowNode. A cognitive element event can be understood as a document-accurately recognized part associated with a MultiRelationNode and the relationship between MultiRelationNodes. The essence of cognitive element combination is the orderly combination of multiple cognitive elements (document-accurately recognized parts) from multiple categories, thereby obtaining the combined product MultiRelationNode.ActionBusiness. Quantitative change then leads to qualitative change. If more categories of cognitive relationship information are accumulated and they meet the predefined requirements, an increasingly accurate level category MultiRelationNode.ActionBusinessType will be generated.

[0640] Third, a document-level category change impacts the associated impact system. A change in the category of a first MultiRelationNode, determined by a high-confidence classification method, affects the business-level category of a second MultiRelationNode or a second document with low confidence, based on the relationship between high-level continuation relationships (relationship degree, relationship direction, etc.). The scope of documents affected by this category change is controlled by the degree of relationship. This high-confidence cognitive-level category transfer utilizes cognitive relationship information and the non-document individual attributes to modify the high-level continuation relationships and the relationships between high-level continuation relationships corresponding to the cognitive element. This includes: updating the document entity flow relationship or document content flow relationship corresponding to the document operated by the cognitive element based on the cognitive relationship information, and updating the relationship between the document entity flow relationship or document content flow relationship corresponding to the document operated by the cognitive element and the document entity flow relationship or document content flow relationship corresponding to documents operated by other cognitive elements based on the non-document individual attributes; or updating the folder cognition corresponding to the folder attribute of the document operated by the cognitive element based on the folder attribute, and updating the relationship between the folder cognition corresponding to the folder attribute and the folder cognition corresponding to other folder attributes based on the cognitive relationship information, wherein the non-document individual attributes include the folder attribute.

[0641] The multi-category cognitive element combination system designed in this application is a positive feedback system. The positive feedback is mainly reflected in the fact that users generate cognitive elements of different categories by performing operations on different categories of documents.

[0642] 1) It doesn't matter if document recognition is inaccurate now, as long as it is accurate in the future. For example, if a MultiRelationNode is determined to be composed of multiple categories of cognitive elements, the more fluid the document, the more cognitive element relationship categories it has, and the more accurate the recognition. For example, when a document is just created and there is no collaboration, there are not many accumulated cognitive element relationship categories, only creation category cognitive elements and derivative category cognitive elements, so the FileFlowNode recognition is not very accurate. With editing and collaboration with other users, it will become more accurate. For example, a FileFlowNode that has entered the multi-user collaboration stage has its cognitive element category added with the DataInMotion category, and its recognition accuracy is significantly improved compared to the FileFlowNode that has not collaborated in the past. That is, from the perspective of a FileFlowNode, the fluid document cognition has more cognitive attributes than the non-fluid single document.

[0643] 2) It doesn't matter if the document itself is not accurately recognized, as long as the associated recognition is accurate. This application establishes relationships between high-level continuation relationships. A single, uncoordinated document can be accurately identified through other related high-level continuation relationships, thereby associating documents that are not accurately recognized. For example, some documents downloaded from the Internet in order to edit a document belong to the auxiliary document family FileFlowNode and will be affected by the update of the level category of the most closely related main document family FileFlowNode.

[0644] 3) This application has designed a mechanism for highly reliable document recognition results to influence the classification of multiple related documents. For example, the more documents are manually labeled, the more accurate a large number of related documents will be. For example, if this system is connected to a document collaboration system such as OA, network disk, SharePoint, ERP, etc., uploaded collaborative documents will receive relatively accurate recognition, thereby further improving the recognition accuracy of multiple closely related documents.

[0645] This solves the problem of relationship management of documents on PCs in actual application scenarios.

[0646] The above describes in detail the method provided by the embodiment of the present application in conjunction with Figures 1 to 8. The following describes in detail the embodiment of the device of the present application in conjunction with Figures 9 and 10. It should be understood that the description of the method embodiment corresponds to the description of the device embodiment. Therefore, for parts not described in detail, reference can be made to the above method embodiment.

[0647] Figure 9 is a schematic block diagram of a device 900 for obtaining document cognition provided in an embodiment of the present application. The device 900 can be implemented by software, hardware, or a combination of the two. The device 900 provided in an embodiment of the present application can implement the method flow shown in the embodiment of the present application, and the device 900 includes: an acquisition module 910 and a processing module 920, wherein the acquisition module 910 is used to obtain a basic event; the processing module 920 determines a cognitive element based on the basic event and / or a combination of the basic events, and the cognitive element includes cognitive relationship information determined by the basic event and cognitive attribute information determined by the basic event; determines at least one high-level continuation relationship related to the cognitive element based on the cognitive relationship information and the cognitive attribute information; and selectively updates the at least one high-level continuation relationship and / or the relationship between high-level continuation relationships based on the at least one high-level continuation relationship.

[0648] Optionally, the cognitive element further includes addressing data between the document and the high-level continuation relationship.

[0649] Optionally, the processing module 920 is further used to determine the cognitive attribute information based on at least two of the following information: application attributes, device attributes, user attributes, path attributes, document extension attributes, and time attributes of the basic event and / or the combination of the basic events.

[0650] Optionally, the processing module 920 is also used to determine the cognitive relationship information based on at least one of the following information: intra-entity relationships of individual document attributes, inter-entity relationships of individual document attributes, and the inter-entity relationships of individual document attributes include document mirror entity derivation relationships and document network transmission relationships.

[0651] Optionally, the processing module 920 is further used to store the addressing data between the document and the advanced continuation relationship according to at least one of the following information: document extended metadata stores the addressing data between the document and the advanced continuation relationship, a predefined database or file stores the addressing data between document attributes and the advanced continuation relationship.

[0652] Optionally, the processing module 920 is also used to determine the addressing data method for storing the document and the advanced continuation relationship in the document extension attribute based on the location attribute of the destination document when there is a risk of failing to accumulate the cognition relationship information between the source document and the advanced continuation relationship according to the cognitive relationship information determined by the basic event.

[0653] Optionally, the processing module 920 is also used to determine the cognitive element of the entity relationship of the individual attribute of the document based on the basic event and / or the combination of the basic events, update the cognition of the source document mirror or update the cognitive attribute information of the high-level continuation relationship corresponding to the source document.

[0654] Optionally, the processing module 920 is also used to determine the cognitive elements of the entity relationship of the individual attributes of the document based on the basic event and / or the combination of the basic events, create the cognition of the new document mirror, and\or update the cognitive attribute information of the high-level continuation relationship corresponding to the source document.

[0655] Optionally, the cognition of the created new document image and the cognition of the source document image are determined based on a decision by combining the cognition of the source document image and the cognition element.

[0656] Optionally, the processing module 920 is specifically used to: determine the cognitive element of the document mirror entity derivative relationship based on the basic event and / or the combination of the basic events, and based on the decision, selectively update the cognitive attribute information of the at least one high-level continuation relationship.

[0657] Optionally, the processing module 920 is specifically used to: determine the cognitive element of the document network transmission relationship based on the basic event and / or the combination of the basic events, and based on the decision, selectively update the cognitive attribute information of the at least one high-level continuation relationship.

[0658] Optionally, the processing module 920 is further configured to update a class of high-level continuation relationships based on the cognitive elements of the document mirror entity derived relationship and the cognitive elements of the document network transmitted relationship.

[0659] Optionally, the processing module 920 is specifically configured to simultaneously update at least one of the high-level continuation relationships and the relationship between the high-level continuation relationships, where the relationship between the high-level continuation relationships includes a degree of relationship and / or a direction of relationship.

[0660] Optionally, the processing module 920 is further configured to create a new category cognition element according to the high-level continuation relationship, maintain the order of the category cognition elements according to the high-level continuation relationship, and update the high-level continuation relationship determined by the addressing data according to the addressing data of the high-level continuation relationship.

[0661] Optionally, the processing module 920 is further configured to update the high-level continuation relationship determined by the addressing data according to the addressing data of the high-level continuation relationship if the cognize element is any one of the cognize element category combinations predefined for the high-level continuation relationship.

[0662] Optionally, the device 900 is used to perform hierarchical classification on the document.

[0663] Optionally, the processing module 920 is also used to determine that the first high-level continuation relationship is a predefined level category and\or determine that the document corresponding to the first high-level continuation relationship is a predefined level category if the order of subject attributes meets the decision; or the processing module 920 is also used to determine that the first high-level continuation relationship is a predefined level category and\or determine that the document corresponding to the first high-level continuation relationship is a predefined level category if the combination of multiple stored category cognitive elements meets the decision.

[0664] Optionally, the processing module 920 is also used to update the level category of the second high-level continuation relationship and\or the level category of the second document corresponding to the second high-level continuation relationship based on the relationship between high-level continuation relationships if the level category of the first high-level continuation relationship changes, and the method for determining the level category of the first high-level continuation relationship is a predefined level category determination method; the method for determining the level category of the first high-level continuation relationship is a predefined level category determination method, which is determined based on at least one of the following information: the order of the subject attributes of the high-level continuation relationship conforms to the decision, the combination of multiple category cognitive elements stored in the high-level continuation relationship conforms to the decision, manual labeling, and level categories obtained from third-party applications.

[0665] Optionally, the processing module 920 is also used to update the level category of the second high-level continuation relationship and\or the level category of the second document corresponding to the second high-level continuation relationship based on the relationship between high-level continuation relationships if the level category of the first high-level continuation relationship changes, and the method for determining the level category of the first high-level continuation relationship is a predefined level category determination method, wherein at least the first type of cognitive attribute is used in the determination of the level category of the first high-level continuation relationship, and at least the second type of cognitive attribute is used in the determination of the relationship between the high-level continuation relationships.

[0666] Optionally, the processing module 920 is specifically configured to: change the level category of the second document or the second high-level continuation relationship in response to a change in the level category of the first high-level continuation relationship.

[0667] Optionally, the processing module 920 is specifically configured to determine the level category of the second document or the second high-level continuation relationship by comparing and determining the cognitive meta-category combination of the first high-level continuation relationship and the cognitive meta-category combination of the second high-level continuation relationship.

[0668] Optionally, in the step of updating the level category of the second document and / or the second high-level continuation relationship based on the relationship between the high-level continuation relationships, the relationship between the high-level continuation relationships includes the degree of the relationship and / or the direction of the relationship.

[0669] Optionally, the processing module 920 is specifically configured to determine a level category of the second document or the second high-level continuation relationship based on the first high-level continuation relationship.

[0670] Optionally, the processing module 920 is further configured to change the level category of the second document or the second high-level continuation relationship in response to the cognitive element of the first document based on the addressing data and the high-level continuation relationship attributes between the first document and the high-level continuation relationship.

[0671] Optionally, the processing module 920 is further configured to change the level category of the second document or the second high-level continuation relationship in response to the cognitive element of the first document based on the addressing data between the first document and the high-level continuation relationship, the attributes of the high-level continuation relationship, and the degree of relationship between the high-level continuation relationships.

[0672] Optionally, the processing module 920 is further configured to determine a combination of the level category of the document and network data transmitted by the document over the network, and transmit the combination to a predefined network device.

[0673] Optionally, the device 900 is used to generate a document family distribution graph.

[0674] Optionally, the processing module 920 is further configured to aggregate multiple high-level continuation relationships of different categories based on the relationships between the high-level continuation relationships to generate an audit drawing, where the audit drawing is used to represent the distribution of documents on the user device.

[0675] Optionally, the device 900 is used to control access to documents, where document access includes opening a file, writing to a file, deleting a file, changing file permissions, changing file attributes, opening an email message in a mail store, deleting an email in a mail store, retrieving a document from a document management system, storing a document in a document management system, or any act of accessing a document or document repository.

[0676] Optionally, the access or use of the control document is based on a combination of the following three: the first high-level continuation relationship corresponding to the document, the relationship between the first high-level continuation relationship and the second high-level continuation relationship, and the level category of the second high-level continuation relationship.

[0677] The present application also provides a device for obtaining document recognition, and its schematic structure is still shown in FIG9 . The device 900 for obtaining document recognition includes an acquisition module 910 and a processing module 920 .

[0678] An acquisition module 910 is configured to obtain a basic event and / or a combination of basic events;

[0679] Processing module 920 is configured to determine a cognitive element, the cognitive element including cognitive relationship information and cognitive attribute information, wherein the cognitive relationship information is determined based on the basic event and / or a combination of the basic events, and the cognitive relationship information includes one of relationship combination members consisting of inter-entity relationships of various document individual attributes and intra-entity relationships of various document individual attributes, and the cognitive attribute information includes document individual attributes and non-document individual attributes corresponding to the cognitive relationship information;

[0680] The processing module 920 is further configured to update, based on the cognize, a high-level continuation relationship corresponding to the cognize and / or a relationship between high-level continuation relationships, wherein the high-level continuation relationship is a combination of cognizes associated with a corresponding data set and / or data updated by the combination of cognizes based on a rule;

[0681] The processing module 920 is further configured to update the level category of the document corresponding to the updated high-level continuation relationship when the updated high-level continuation relationship and / or the relationship between the high-level continuation relationships meets a predefined decision, wherein the predefined decision includes high-level continuation relationship features and their corresponding level categories.

[0682] Optionally, the processing module 920 is specifically used to: update the high-level continuation relationship and / or the relationship between high-level continuation relationships corresponding to the cognitive element according to the type of the cognitive element, wherein the type of the cognitive element is determined according to the cognitive relationship information and / or the creation information of the high-level continuation relationship corresponding to the cognitive element, and the creation information indicates that the high-level continuation relationship corresponding to the cognitive element has been created or the high-level continuation relationship corresponding to the cognitive element has not been created.

[0683] Optionally, the type of the cognitive element meets the requirement of creating a member cognitive element of a high-level continuation relationship, and the processing module 920 is specifically used to: create a high-level continuation relationship corresponding to the cognitive element.

[0684] Optionally, the type of the cognitive element conforms to the continuation member cognitive element of the high-level continuation relationship, and the processing module 920 is specifically used to: continue the high-level continuation relationship corresponding to the cognitive element.

[0685] Optionally, the high-level continuation relationship includes at least one of the following relationships: document mirror recognition, document entity flow relationship, document content flow relationship, and folder recognition.

[0686] Optionally, the processing module 920 is specifically configured to update the folder cognition corresponding to the folder attribute based on the folder attribute of the document of the cognitive meta-operation, wherein the folder cognition is a type of the high-level continuation relationship.

[0687] Optionally, the cognitive element is a cognitive element that expresses the document mirror entity derivation relationship or a cognitive element that expresses the document network transmission relationship. The processing module 920 is specifically used to: associate the document operated by the cognitive element and the copy generated by the cognitive element with the same document entity flow relationship, where the document entity flow relationship is a type of the high-level continuation relationship.

[0688] Optionally, the document of the cognitive meta-operation has a unique corresponding document mirror cognition, and the document mirror cognition is a type of the high-level continuation relationship.

[0689] Optionally, the processing module 920 is specifically used to: based on the storage location of the document of the cognitive meta-operation, select at least one of the following methods to store the extended attributes of the document of the cognitive meta-operation, and the extended attributes include addressing data of high-level continuation relationships: embedding the extended attributes into the extensible attributes in the file format; storing the extended attributes for modifying the document content; encrypting and encapsulating the extended attributes and the main text file into one file; storing the extended attributes in the extensible attributes part of the file system; storing the extended attributes in a predefined database or file.

[0690] Optionally, the processing module 920 is specifically configured to: when the cognitive element indicates that the destination document is at risk of losing the addressing data of the high-level continuation relationship of the source document, ensure that the addressing data of the high-level continuation relationship of the destination document is not lost.

[0691] Optionally, when executing the step of updating the high-level continuation relationship corresponding to the cognitive element and / or the relationship between high-level continuation relationships, the processing module 920 uses at least the individual identifier stored in the document extension attribute, and the individual identifier is used to determine the correspondence between the document and the high-level continuation relationship.

[0692] Optionally, the individual identifier includes at least one of the following: a document identifier, a document mirror identifier, and an identifier that enables a one-to-one mapping of a document to a document mirror.

[0693] Optionally, the relationship between the high-level continuation relationships includes the degree of the relationship between the high-level continuation relationships and the direction of the relationship between the high-level continuation relationships, wherein the direction of the relationship between the high-level continuation relationships is determined based on the source and destination of the cognitive relationship information.

[0694] Optionally, the processing module 920 is specifically configured to modify the high-level continuation relationship corresponding to the cognitive element and the relationship ...

Claims

1. A method for obtaining document recognition, characterized in that: The method comprises: Obtaining a base event and / or a combination of base events; Determine a cognitive element, wherein the cognitive element includes cognitive relationship information and cognitive attribute information, wherein the cognitive relationship information is determined according to the basic event and / or a combination of the basic events, and the cognitive relationship information includes one of relationship combination members constituted by inter-entity relationships of various document individual attributes and intra-entity relationships of various document individual attributes, and the cognitive attribute information includes document individual attributes and non-document individual attributes corresponding to the cognitive relationship information; Based on the cognitive element, updating the high-level continuation relationship and / or the relationship between high-level continuation relationships corresponding to the cognitive element, wherein the high-level continuation relationship is the data updated by the cognitive element combination associated with the corresponding data set and / or the cognitive element combination based on the rule; When the updated high-level continuation relationship and / or the relationship between high-level continuation relationships conforms to a predefined decision, the level category of the document corresponding to the updated high-level continuation relationship is updated, wherein the predefined decision includes high-level continuation relationship features and their corresponding level categories.

2. The method according to claim 1, characterized in that: The updating of the high-level continuation relationship and / or the relationship between high-level continuation relationships corresponding to the cognize element based on the cognize element includes: The high-level continuation relationship and / or the relationship between high-level continuation relationships corresponding to the cognitive element are updated according to the type of the cognitive element, wherein the type of the cognitive element is determined based on the cognitive relationship information and / or the creation information of the high-level continuation relationship corresponding to the cognitive element, and the creation information indicates that the high-level continuation relationship corresponding to the cognitive element has been created or the high-level continuation relationship corresponding to the cognitive element has not been created.

3. The method according to claim 2, characterized in that The type of the cognizer conforms to the creation member cognizer of the high-level continuation relationship, and the updating of the high-level continuation relationship corresponding to the cognizer according to the type of the cognizer includes: A high-level continuation relationship corresponding to the cognitive element is newly created.

4. The method according to claim 2, characterized in that: The type of the cognize element conforms to the continuation member cognize of the high-level continuation relationship, and the updating of the high-level continuation relationship corresponding to the cognize element according to the type of the cognize element includes: Continue the high-level continuation relationship corresponding to the cognitive element.

5. The method according to any one of claims 1 to 4, characterized in that The high-level continuation relationship includes at least one of the following relationships: Document mirror recognition, document entity flow relationship, document content flow relationship, and folder recognition.

6. The method according to any one of claims 1 to 5, characterized in that The updating of the high-level continuation relationship and / or the relationship between high-level continuation relationships corresponding to the cognize element based on the cognize element includes: Based on the folder attribute of the document of the cognitive meta-operation, the folder cognition corresponding to the folder attribute is updated, wherein the folder cognition is a type of the high-level continuation relationship.

7. The method according to any one of claims 1 to 6, characterized in that The cognates are cognates that express document mirror entity derivation relationships or cognates that express document network transmission relationships. The updating of the high-level continuation relationships and / or the relationships between high-level continuation relationships corresponding to the cognates based on the cognates includes: The document operated by the cognitive element and the copy generated by the cognitive element are associated with the same document entity flow relationship, wherein the document entity flow relationship is a type of the high-level continuation relationship.

8. The method according to any one of claims 1 to 7, characterized in that The document of the cognitive meta-operation has a unique corresponding document mirror cognition, and the document mirror cognition is a type of the high-level continuation relationship.

9. The method according to any one of claims 1 to 8, characterized in that The determining of cognitive elements comprises: Based on the storage location of the document of the cognitive meta-operation, at least one of the following methods is selected to store the extended attributes of the document of the cognitive meta-operation, wherein the extended attributes include addressing data of a high-level continuation relationship: Embedding the extended attribute into an extensible attribute in a file format; Used to modify the document content to store the extended attribute; Encrypting and packaging the extended attributes and the text file into one file; Storing the extended attribute in the file system extendable attribute part; The extended attributes are stored in a predefined database or file.

10. The method according to any one of claims 1 to 9, characterized in that The updating of the high-level continuation relationship and / or the relationship between high-level continuation relationships corresponding to the cognize element based on the cognize element includes: When the cognitive element indicates that the destination document is at risk of losing the addressing data of the higher-level continuation relationship of the source document, the higher-level continuation relationship addressing data of the destination document is prevented from being lost.

11. The method according to any one of claims 1 to 10, characterized in that In the step of updating the high-level continuation relationship and / or the relationship between high-level continuation relationships corresponding to the cognitive element, at least the individual identifier stored in the document extended attribute is used, and the individual identifier is used to determine the corresponding relationship between the document and the high-level continuation relationship.

12. The method according to claim 11, characterized in that The individual identifier includes at least one of the following: Document ID and document image ID can map a document to a document image one-to-one.

13. The method according to any one of claims 1 to 12, characterized in that The relationship between the high-level continuation relationships includes the degree of the relationship between the high-level continuation relationships and the direction of the relationship between the high-level continuation relationships, wherein the direction of the relationship between the high-level continuation relationships is determined based on the source and destination of the cognitive relationship information.

14. The method according to any one of claims 1 to 13, characterized in that The updating of the high-level continuation relationship and / or the relationship between high-level continuation relationships corresponding to the cognize element based on the cognize element includes: The high-level continuation relationship corresponding to the cognitive element and the relationship between the high-level continuation relationships are modified based on the cognitive relationship information and the non-document individual attribute.

15. The method according to claim 14, characterized in that The modifying the high-level continuation relationship corresponding to the cognitive element and the relationship between the high-level continuation relationships based on the cognitive relationship information and the non-document individual attribute includes: Update the document entity flow relationship or document content flow relationship corresponding to the document of the cognitive meta-operation based on the cognitive relationship information, and update the relationship between the document entity flow relationship or document content flow relationship corresponding to the document of the cognitive meta-operation and the document entity flow relationship or document content flow relationship corresponding to documents of other cognitive meta-operations based on the non-document individual attributes; or The folder cognition corresponding to the folder attribute of the document based on the cognitive meta-operation is updated, and the relationship between the folder cognition corresponding to the folder attribute and the folder cognition corresponding to other folder attributes is updated based on the cognitive relationship information, wherein the non-document individual attributes include the folder attribute.

16. The method according to any one of claims 1 to 15, characterized in that The method is applied to hierarchically classify the documents.

17. The method according to any one of claims 1 to 16, characterized in that Before updating the level category of the document corresponding to the updated high-level continuation relationship, the method further includes: Determine the attribute characteristics of the high-level continuation relationship, wherein the attribute characteristics include at least one of the following: boundary characteristics, importance of the high-level continuation relationship, folder attributes corresponding to the high-level continuation relationship, subject attribute sequence corresponding to the high-level continuation relationship, and application combination mode of the high-level continuation relationship; When the updated high-level continuation relationship and / or the relationship between the high-level continuation relationships meets the predefined decision, updating the level category of the document corresponding to the updated high-level continuation relationship includes: When the attribute features of the high-level continuation relationship or the combination of the attribute features of the high-level continuation relationship conforms to a first decision subset in a predefined decision, the level category of the document corresponding to the updated high-level continuation relationship is updated, and the first decision subset includes the attribute features of the high-level continuation relationship and its corresponding level category and\or the combination of the attribute features of the high-level continuation relationship and its corresponding level category.

18. The method according to claim 17, characterized in that When the attribute feature of the high-level continuation relationship or the combination of the attribute features of the high-level continuation relationship meets the first decision subset in the predefined decision, updating the level category of the document corresponding to the updated high-level continuation relationship includes: If the boundary feature and the subject attribute sequence of the high-level continuation relationship meet the first decision subset, the high-level continuation relationship meeting the attribute feature combination is determined to be a predefined level category, and / or, the document corresponding to the high-level continuation relationship meeting the attribute feature combination is determined to be a predefined level category, and the subject attribute sequence includes a time series of subject attributes; or If the boundary characteristics and importance of the high-level continuation relationship meet the first decision subset, determine that the high-level continuation relationship that meets the attribute feature combination is a predefined level category, and / or, determine that the document corresponding to the high-level continuation relationship that meets the attribute feature combination is a predefined level category; or If the boundary characteristics and application combination mode of the advanced continuation relationship meet the first decision subset, determining that the advanced continuation relationship meeting the attribute feature combination is a predefined level category, and / or determining that the document corresponding to the advanced continuation relationship meeting the attribute feature combination is a predefined level category; or If the combination of the boundary feature, the application combination mode and the importance of the advanced continuation relationship meets the first decision subset, determining that the advanced continuation relationship meeting the attribute feature combination is a predefined level category, and / or determining that the document corresponding to the advanced continuation relationship meeting the attribute feature combination is a predefined level category; or If the combination of boundary features of the advanced continuation relationship, the application combination method and the folder attributes corresponding to the advanced continuation relationship meets the first decision subset, the advanced continuation relationship that meets the attribute feature combination is determined to be a predefined level category, and\or, the document corresponding to the advanced continuation relationship that meets the attribute feature combination is determined to be a predefined level category.

19. The method according to any one of claims 1 to 18, characterized in that When the updated high-level continuation relationship and / or the relationship between the high-level continuation relationships meets the predefined decision, updating the level category of the document corresponding to the updated high-level continuation relationship includes: When the folder attribute of the updated high-level continuation relationship meets the second decision subset in the predefined decision, the level category of the document corresponding to the updated high-level continuation relationship is updated, and the second decision subset includes the folder attribute of the high-level continuation relationship and its corresponding level category.

20. The method according to any one of claims 1 to 19, characterized in that When the updated high-level continuation relationship and / or the relationship between the high-level continuation relationships meets the predefined decision, updating the level category of the document corresponding to the updated high-level continuation relationship includes: When the relationship between the first high-level continuation relationship and the second high-level continuation relationship, the attribute characteristics of the first high-level continuation relationship, and the combination of the level categories of the second high-level continuation relationship meet the third decision subset in the predefined decision, the level category of the document corresponding to the first high-level continuation relationship is updated, the first high-level continuation relationship is the high-level continuation relationship corresponding to the document of the cognitive meta-operation, and the third decision subset includes the attribute characteristics of the first high-level continuation relationship, the relationship between the first high-level continuation relationship and the second high-level continuation relationship, the combination of the level categories of the second high-level continuation relationship and their corresponding level categories.

21. The method according to any one of claims 1 to 20, characterized in that The cognates are entity relationship category cognates of individual document attributes, and updating the high-level continuation relationship and / or the relationship between high-level continuation relationships corresponding to the cognates based on the cognates includes: A high-level continuation relationship corresponding to the destination document is newly created, wherein at least part of the cognitive attributes of the high-level continuation relationship corresponding to the destination document comes from the cognitive attributes of the high-level continuation relationship corresponding to the source document.

22. The method according to any one of claims 1 to 21, characterized in that The method further comprises: In response to a change in the level category of a first high-level continuation relationship, at least one of the following methods is used to determine whether to change the level category of a third high-level continuation relationship or the level category of a document corresponding to the third high-level continuation relationship, wherein the first high-level continuation relationship is a high-level continuation relationship corresponding to the document of the cognitive meta-operation: comparing the credibility of the method for determining the level category of the first high-level continuation relationship and the method for determining the level category of the third high-level continuation relationship; determining whether the credibility of the method for determining the level category of the first high-level continuation relationship is greater than a threshold; determining whether a method for determining the level category of the first high-level continuation relationship is a predefined determination method; comparing directions of the relationships between the first high-level continuation relationship and the third high-level continuation relationship; The importance of the first high-level continuation relationship and the importance of the third high-level continuation relationship are compared.

23. The method according to any one of claims 1 to 22, characterized in that The updating of the high-level continuation relationship and / or the relationship between high-level continuation relationships corresponding to the cognize element based on the cognize element includes: In response to the document of the cognitive meta-operation, the high-level continuation relationship corresponding to the document is determined based on the addressing data of the high-level continuation relationship stored in the extended attribute of the document.

24. The method according to any one of claims 1 to 23, characterized in that The updating of the high-level continuation relationship and / or the relationship between high-level continuation relationships corresponding to the cognize element based on the cognize element includes: A high-level continuation relationship corresponding to the cognitive element and a relationship between the newly created high-level continuation relationship and the existing high-level continuation relationship are created.

25. The method according to any one of claims 1 to 24, characterized in that The method further comprises: The document and / or the level category of the document is transmitted to a predefined network device according to the level category of the document and the network data of the document transmitted by the network.

26. The method according to any one of claims 1 to 25, characterized in that The method further comprises: Based on the relationship between the advanced continuation relationships, a plurality of advanced continuation relationships of different categories are collected to generate an audit drawing. The graph is used to represent the distribution of documents on user devices.

27. The method according to any one of claims 1 to 26, characterized in that The method further comprises: Based on a plurality of high-level continuation relationships, an audit drawing is generated, wherein the audit drawing is used to represent a description of a working situation of subject attributes.

28. The method according to any one of claims 1 to 27, characterized in that The method further comprises: Based on at least one high-level continuation relationship and a predefined knowledge determination model or a data-driven model, it is determined whether the user behavior is normal.

29. The method according to any one of claims 1 to 28, characterized in that The method is applied to control the access to a document, wherein the access to the document includes: opening a file, writing to a file, deleting a file, changing file permissions, changing file attributes, opening an email message in a mail storage, deleting an email in a mail storage, retrieving a document from a document management system, synchronizing with a predefined document of a document management system, storing a document in a document management system, or any behavior of accessing a document or a document repository.

30. The method according to claim 29, characterized in that During the access process of the document, the level category of the document, the storage location of the application corresponding to the document, and the individual identifier in the document extension attribute are used.

31. The method according to claim 29 or 30, characterized in that The method further comprises: Access or use of the document is controlled based on a combination of the following three: a first high-level continuation relationship corresponding to the document, a relationship between the first high-level continuation relationship and a fourth high-level continuation relationship, and a level category of the fourth high-level continuation relationship.

32. The method according to claim 31, characterized in that The method for determining the level category of the fourth high-level continuation relationship is a predefined level category determination method or the fourth high-level continuation relationship is a combination of predefined cognitive meta-categories.

33. A device for obtaining document recognition, characterized in that: The device comprises: An acquisition module, used for acquiring a basic event and / or a combination of basic events; A processing module, used to determine a cognitive element, wherein the cognitive element includes cognitive relationship information and cognitive attribute information, wherein the cognitive relationship information is determined according to the basic event and / or the combination of the basic events, and the cognitive relationship information includes one of the relationship combination members constituted by the inter-entity relationship of each type of document individual attribute and the intra-entity relationship of each type of document individual attribute, and the cognitive attribute information includes the document individual attribute and the non-document individual attribute corresponding to the cognitive relationship information; The processing module is further used to update the high-level continuation relationship and / or the relationship between high-level continuation relationships corresponding to the cognitive element based on the cognitive element, wherein the high-level continuation relationship is data updated based on the rule by the cognitive element combination associated with the corresponding data set and / or the combination of multiple types of cognitive elements; The processing module is further configured to update the level category of the document corresponding to the updated high-level continuation relationship when the updated high-level continuation relationship and / or the relationship between the high-level continuation relationships conforms to a predefined decision, wherein the predefined decision includes high-level continuation relationship features and their corresponding level categories.

34. The device according to claim 33, characterized in that The processing module is specifically used for: The high-level continuation relationship and / or the relationship between high-level continuation relationships corresponding to the cognitive element are updated according to the type of the cognitive element, wherein the type of the cognitive element is determined based on the cognitive relationship information and / or the creation information of the high-level continuation relationship corresponding to the cognitive element, and the creation information indicates that the high-level continuation relationship corresponding to the cognitive element has been created or the high-level continuation relationship corresponding to the cognitive element has not been created.

35. The device according to claim 34, characterized in that The type of the cognition element conforms to the creation member cognition element of the high-level continuation relationship, and the processing module is specifically used for: A high-level continuation relationship corresponding to the cognitive element is newly created.

36. The device according to claim 34, characterized in that The type of the cognate element conforms to the continuation member cognate element of the high-level continuation relationship, and the processing module is specifically used for: Continue the high-level continuation relationship corresponding to the cognitive element.

37. The device according to any one of claims 33 to 36, characterized in that The high-level continuation relationship includes at least one of the following relationships: Document mirror recognition, document entity flow relationship, document content flow relationship, and folder recognition.

38. The device according to any one of claims 33 to 37, characterized in that The processing module is specifically used for: Based on the folder attribute of the document of the cognitive meta-operation, the folder cognition corresponding to the folder attribute is updated, wherein the folder cognition is a type of the high-level continuation relationship.

39. The device according to any one of claims 33 to 38, characterized in that The cognitive element is a cognitive element that expresses the derived relationship of the document mirror entity or a cognitive element that expresses the relationship of the document being transmitted by the network. The processing module is specifically used for: The document operated by the cognitive element and the copy generated by the cognitive element are associated with the same document entity flow relationship, wherein the The document entity flow relationship is a type of the advanced continuation relationship.

40. The device according to any one of claims 33 to 39, characterized in that The document of the cognitive meta-operation has a unique corresponding document mirror cognition, and the document mirror cognition is a type of the high-level continuation relationship.

41. The device according to any one of claims 33 to 40, characterized in that The processing module is specifically used for: Based on the storage location of the document of the cognitive meta-operation, at least one of the following methods is selected to store the extended attributes of the document of the cognitive meta-operation, wherein the extended attributes include addressing data of a high-level continuation relationship: Embedding the extended attribute into an extensible attribute in a file format; Used to modify the document content to store the extended attribute; Encrypting and packaging the extended attributes and the text file into one file; Storing the extended attribute in the file system extendable attribute part; The extended attributes are stored in a predefined database or file.

42. The device according to any one of claims 33 to 41, characterized in that The processing module is specifically used for: When the cognitive element indicates that the destination document is at risk of losing the addressing data of the higher-level continuation relationship of the source document, the higher-level continuation relationship addressing data of the destination document is prevented from being lost.

43. The device according to any one of claims 33 to 42, characterized in that The processing module uses at least the individual identifier stored in the document extension attribute when executing the step of updating the high-level continuation relationship and / or the relationship between high-level continuation relationships corresponding to the cognitive element, and the individual identifier is used to determine the correspondence between the document and the high-level continuation relationship.

44. The device according to claim 43, characterized in that The individual identifier includes at least one of the following: Document ID and document image ID can map a document to a document image one-to-one.

45. The device according to any one of claims 33 to 44, characterized in that The relationship between the high-level continuation relationships includes the degree of the relationship between the high-level continuation relationships and the direction of the relationship between the high-level continuation relationships, wherein the direction of the relationship between the high-level continuation relationships is determined based on the source and destination of the cognitive relationship information.

46. ​​The device according to any one of claims 33 to 45, characterized in that The processing module is specifically used for: The high-level continuation relationship corresponding to the cognitive element and the relationship between the high-level continuation relationships are modified based on the cognitive relationship information and the non-document individual attribute.

47. The device according to claim 46, characterized in that The processing module is specifically used for: Update the document entity flow relationship or document content flow relationship corresponding to the document of the cognitive meta-operation based on the cognitive relationship information, and update the relationship between the document entity flow relationship or document content flow relationship corresponding to the document of the cognitive meta-operation and the document entity flow relationship or document content flow relationship corresponding to the documents of other cognitive meta-operations based on the non-document individual attributes; or The folder cognition corresponding to the folder attribute of the document based on the cognitive meta-operation is updated, and the relationship between the folder cognition corresponding to the folder attribute and the folder cognition corresponding to other folder attributes is updated based on the cognitive relationship information, wherein the non-document individual attributes include the folder attribute.

48. The device according to any one of claims 33 to 47, characterized in that The device is used to classify the documents in a hierarchical manner.

49. The device according to any one of claims 33 to 48, characterized in that The processing module is also used for: Determine the attribute characteristics of the high-level continuation relationship, wherein the attribute characteristics include at least one of the following: boundary characteristics, importance of the high-level continuation relationship, folder attributes corresponding to the high-level continuation relationship, subject attribute sequence corresponding to the high-level continuation relationship, and application combination mode of the high-level continuation relationship; Wherein, the processing module is specifically used for: When the attribute features of the high-level continuation relationship or the combination of the attribute features of the high-level continuation relationship conforms to a first decision subset in a predefined decision, the level category of the document corresponding to the updated high-level continuation relationship is updated, and the first decision subset includes the attribute features of the high-level continuation relationship and its corresponding level category and\or the combination of the attribute features of the high-level continuation relationship and its corresponding level category.

50. The device according to claim 49, characterized in that The processing module is specifically used for: If the boundary feature and the subject attribute sequence of the high-level continuation relationship meet the first decision subset, the high-level continuation relationship meeting the attribute feature combination is determined to be a predefined level category, and / or, the document corresponding to the high-level continuation relationship meeting the attribute feature combination is determined to be a predefined level category, and the subject attribute sequence includes a time series of subject attributes; or If the boundary characteristics and importance of the high-level continuation relationship meet the first decision subset, determine that the high-level continuation relationship that meets the attribute feature combination is a predefined level category, and / or, determine that the document corresponding to the high-level continuation relationship that meets the attribute feature combination is a predefined level category; or If the boundary characteristics and application combination mode of the advanced continuation relationship meet the first decision subset, determining that the advanced continuation relationship meeting the attribute feature combination is a predefined level category, and / or determining that the document corresponding to the advanced continuation relationship meeting the attribute feature combination is a predefined level category; or If the combination of the boundary feature, the application combination mode and the importance of the advanced continuation relationship meets the first decision subset, determining that the advanced continuation relationship meeting the attribute feature combination is a predefined level category, and / or determining that the document corresponding to the advanced continuation relationship meeting the attribute feature combination is a predefined level category; or If the combination of boundary features of the advanced continuation relationship, the application combination method and the folder attributes corresponding to the advanced continuation relationship meets the first decision subset, the advanced continuation relationship that meets the attribute feature combination is determined to be a predefined level category, and\or, the document corresponding to the advanced continuation relationship that meets the attribute feature combination is determined to be a predefined level category.

51. The device according to any one of claims 33 to 50, characterized in that The processing module is specifically used for: When the folder attribute of the updated high-level continuation relationship meets the second decision subset in the predefined decision, the level category of the document corresponding to the updated high-level continuation relationship is updated, and the second decision subset includes the folder attribute of the high-level continuation relationship and its corresponding level category.

52. The device according to any one of claims 33 to 51, characterized in that The processing module is specifically used for: When the relationship between the first high-level continuation relationship and the second high-level continuation relationship, the attribute characteristics of the first high-level continuation relationship, and the combination of the level categories of the second high-level continuation relationship meet the third decision subset in the predefined decision, the level category of the document corresponding to the first high-level continuation relationship is updated, the first high-level continuation relationship is the high-level continuation relationship corresponding to the document of the cognitive meta-operation, and the third decision subset includes the attribute characteristics of the first high-level continuation relationship, the relationship between the first high-level continuation relationship and the second high-level continuation relationship, the combination of the level categories of the second high-level continuation relationship and their corresponding level categories.

53. The device according to any one of claims 33 to 52, characterized in that The cognitive element is a cognitive element of the entity relationship category of the document individual attribute, and the processing module is specifically used for: A high-level continuation relationship corresponding to the destination document is newly created, wherein at least part of the cognitive attributes of the high-level continuation relationship corresponding to the destination document comes from the cognitive attributes of the high-level continuation relationship corresponding to the source document.

54. The device according to any one of claims 33 to 53, characterized in that The processing module is also used for: In response to a change in the level category of a first high-level continuation relationship, at least one of the following methods is used to determine whether to change the level category of a third high-level continuation relationship or the level category of a document corresponding to the third high-level continuation relationship, wherein the first high-level continuation relationship is a high-level continuation relationship corresponding to the document of the cognitive meta-operation: comparing the credibility of the method for determining the level category of the first high-level continuation relationship and the method for determining the level category of the third high-level continuation relationship; determining whether the credibility of the method for determining the level category of the first high-level continuation relationship is greater than a threshold; determining whether a method for determining the level category of the first high-level continuation relationship is a predefined determination method; comparing directions of the relationships between the first high-level continuation relationship and the third high-level continuation relationship; The importance of the first high-level continuation relationship and the importance of the third high-level continuation relationship are compared.

55. The device according to any one of claims 33 to 54, characterized in that The processing module is specifically used for: In response to the document of the cognitive meta-operation, the high-level continuation relationship corresponding to the document is determined based on the addressing data of the high-level continuation relationship stored in the extended attribute of the document.

56. The device according to any one of claims 33 to 55, characterized in that The processing module is specifically used for: A high-level continuation relationship corresponding to the cognitive element and a relationship between the newly created high-level continuation relationship and the existing high-level continuation relationship are created.

57. The device according to any one of claims 33 to 56, characterized in that The processing module is also used for: The document and / or the level category of the document is transmitted to a predefined network device according to the level category of the document and the network data of the document transmitted by the network.

58. The device according to any one of claims 33 to 57, characterized in that The processing module is also used for: Based on the relationship between the high-level continuation relationships, a plurality of high-level continuation relationships of different categories are collected to generate an audit drawing, where the audit drawing is used to represent the distribution of documents on a user device.

59. The device according to any one of claims 33 to 58, characterized in that The processing module is also used for: Based on a plurality of high-level continuation relationships, an audit drawing is generated, wherein the audit drawing is used to represent a description of a working situation of subject attributes.

60. The device according to any one of claims 33 to 59, characterized in that The processing module is also used for: Based on at least one high-level continuation relationship and a predefined knowledge determination model or a data-driven model, it is determined whether the user behavior is normal.

61. The device according to any one of claims 33 to 60, characterized in that The device is used to control access to documents, and the access to the document includes: opening a file, writing to a file, deleting a file, changing file permissions, changing file attributes, opening an email message in a mail storage, deleting an email in the mail storage, retrieving a document from a document management system, synchronizing with a predefined document of a document management system, storing a document in a document management system, or any behavior of accessing a document or a document repository.

62. The device according to claim 61, characterized in that During the access process of the document, the level category of the document, the storage location of the application corresponding to the document, and the individual identifier in the document extension attribute are used.

63. The device according to claim 61 or 62, characterized in that The processing module is also used for: Access or use of the document is controlled based on a combination of the following three: a first high-level continuation relationship corresponding to the document, a relationship between the first high-level continuation relationship and a fourth high-level continuation relationship, and a level category of the fourth high-level continuation relationship.

64. The device according to claim 63, characterized in that The method for determining the level category of the fourth high-level continuation relationship is a predefined level category determination method or the fourth high-level continuation relationship is a combination of predefined cognitive meta-categories.

65. A computing device, characterized in that The device comprises a processor and a memory, wherein the processor is configured to execute instructions stored in the memory so that the computing device performs the method according to any one of claims 1 to 32.

66. A computer-readable storage medium, characterized in that Comprising computer program instructions which, when executed by a computing device, cause the computing device to perform the method of any one of claims 1 to 32.