Data processing method and device based on multi-source data, equipment and storage medium

By obtaining related data from a second data source and performing consistency checks, the problem of lack of office system proofreading in the annotation method is solved, realizing the accuracy of document processing and data consistency across multiple data sources, and is suitable for data processing in multi-data source collaborative scenarios.

CN121808267APending Publication Date: 2026-04-07BEIJING KINGSOFT OFFICE SOFTWARE INC +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In existing technologies, annotation methods lack verification of relevant office data in the office system during document processing, resulting in insufficient comprehensiveness and accuracy of annotation suggestions, which affects the accuracy of document processing.

Method used

By acquiring target data from the first data source and related data from the second data source, consistency checks are performed to accurately locate anomalies and achieve cross-validation and verification of multi-source data.

Benefits of technology

It improves the accuracy and comprehensiveness of data processing, ensures data consistency across multiple data sources in the office system, provides reliable evidence for anomaly identification, adapts to multi-data source collaboration scenarios, and enables effective verification of key information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121808267A_ABST
    Figure CN121808267A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device based on multi-source data, equipment and a storage medium, and relates to the technical field of document operation software. The method comprises the steps of obtaining target data of a first data source; acquiring associated data of the target data from the second data source; and performing consistency detection on the associated data and the target data, and determining an abnormal problem of the target data. Through the above technical means, consistency detection can be performed on the target data of the first data source based on the associated data of the second data source so as to accurately locate the abnormal problem of the first data source, the problem of lack of proofreading of related office data in an office system in the prior art is solved, and the data processing precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data analysis, and in particular to a data processing method and device based on multi-source data, equipment and a storage medium. BACKGROUND

[0002] In the process of daily office work, documents, as the core carrier of information transmission and collaborative work, are widely used in product demand review, project planning, legal contract review, knowledge base maintenance and many other scenarios. With the deepening of enterprise digital transformation, the number of office documents continues to grow, and the relevance between documents and various office systems is increasingly complex, which puts higher requirements on the accuracy, consistency and timeliness of document content. Therefore, during the document processing process, staff often need to edit or modify the document content in a timely manner according to the latest information of the office system to ensure that the document content meets the requirements of the office system.

[0003] In the prior art, in order to improve the efficiency of document processing, corresponding annotation suggestions are added to the document through annotation, so that the user can modify the document content based on the annotation suggestions. However, the annotation method is limited to logical analysis of the document content and giving annotation suggestions, and lacks correction of relevant office data in the office system, resulting in insufficient comprehensiveness and accuracy of the annotation suggestions, thereby affecting the accuracy of document processing. SUMMARY

[0004] The present application provides a data processing method and device based on multi-source data, equipment and a storage medium, which detects the consistency of target data of a first data source based on associated data of a second data source, accurately locates abnormal problems of the first data source, and solves the problem of lack of correction of relevant office data in the office system in the prior art, thereby improving the accuracy of data source processing.

[0005] In a first aspect, the present application provides a data processing method based on multi-source data, comprising: obtaining target data of a first data source; obtaining associated data of the target data from a second data source; detecting the consistency of the associated data and the target data, and determining abnormal problems of the target data.

[0006] In a second aspect, the present application provides a data processing device based on multi-source data, comprising: a target data obtaining module configured to obtain target data of a first data source; an associated data obtaining module configured to obtain associated data of the target data from a second data source; An abnormality problem determination module configured to perform consistency detection on the target data and the associated data, and determine an abnormality problem of the target data.

[0007] In a third aspect, the present application provides a data processing device based on multi-source data, comprising: one or more processors; a memory storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the data processing method based on multi-source data as described in the first aspect.

[0008] In a fourth aspect, the present application provides a storage medium containing computer executable instructions for executing the data processing method based on multi-source data as described in the first aspect when executed by a computer processor.

[0009] In the present application, by obtaining the target data of the first data source as the content to be audited, obtaining the associated data of the target data in the second data source to introduce the calibration support of multi-source data, comparing the consistency of the associated data and the target data to accurately locate the abnormality problem of the target data, cross-checking of multi-source data is realized, the problem of lack of proofreading of office data in the prior art is solved, the comprehensiveness and accuracy of abnormality identification are improved from the data source level, a reliable basis is provided for subsequent abnormality processing, the accuracy of data processing is improved, and the method is suitable for multi-data source collaborative scenarios, effectively checks the key information in the collaborative scenarios, and guarantees the data consistency of multi-data sources. BRIEF DESCRIPTION OF DRAWINGS

[0010] Figure 1 is a flowchart of a data processing method based on multi-source data provided by an embodiment of the present application; Figure 2 is one of the schematic diagrams of the content display interface of the first data source provided by an embodiment of the present application; Figure 3 is one of the schematic diagrams of the content display interface of the first data source provided by an embodiment of the present application; Figure 4 is one of the schematic diagrams of the content display interface of the first data source provided by an embodiment of the present application; Figure 5 is one of the schematic diagrams of the content display interface of the first data source provided by an embodiment of the present application; Figure 6 is one of the schematic diagrams of the space display interface of the team space provided by an embodiment of the present application; Figure 7 is one of the schematic diagrams of the space display interface of the team space provided by an embodiment of the present application; Figure 8Fig. 5 is a schematic view of a content display interface of a first data source according to an embodiment of the present application; Figure 9 Fig. 3 is a schematic view of a space display interface of a team space according to an embodiment of the present application; Figure 10 Fig. 1 is a schematic view of a structure of a data processing device based on multi-source data according to an embodiment of the present application; Figure 11 Fig. 1 is a schematic view of a structure of a data processing device based on multi-source data according to an embodiment of the present application. DETAILED DESCRIPTION

[0011] To make the objects, technical solutions and advantages of the present application clearer, the specific embodiments of the present application are further described in detail below with reference to the drawings. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application. In addition, it should be noted that, for the convenience of description, only parts related to the present application are shown in the drawings, but not all. Before discussing the example embodiments in more detail, it should be mentioned that some example embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be implemented in parallel, concurrently or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when the operations are completed, but can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0012] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be exchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally a class, and do not limit the number of objects, for example, the first object can be one or more. In addition, the specification and claims "and / or" means at least one of the connected objects, and the character " / ", generally indicates that the associated objects before and after are in a "or" relationship.

[0013] In some implementations, the document is a core carrier of the office scenario, and the staff will edit or modify the document according to the latest information of the office to ensure that the content of the document meets the office requirements. In order to improve the efficiency of document processing, the corresponding annotation suggestions are added to the document through the annotation method, so that the user modifies the document content based on the annotation suggestions. In another implementation, intelligent annotation can also be used. The intelligent annotation method analyzes the semantics of the document content, identifies the logical conflicts and semantic inconsistencies of the content, and gives corresponding modification suggestions for the problem content. The problem type and modification suggestion of the problem content are set as the annotation of the problem content in the document. However, even the intelligent annotation method is limited to analyzing the logical conflicts and semantic inconsistencies of the document content and giving annotation suggestions, and lacks proofreading of related office data in the office system, resulting in insufficient comprehensiveness and accuracy of the annotation suggestions, thereby affecting the accuracy of document processing. To solve the above problems, the embodiment provides a data processing method based on multi-source data, to obtain a first data source and a second data source from documents, emails, knowledge bases, calendars, meeting minutes, tasks, and team spaces of an office system, and to detect the consistency of target data of the first data source based on the associated data of the second data source, to accurately locate the abnormal problems of the first data source, so as to subsequently process the abnormal problems of the first data source and ensure the processing accuracy of the data source. Therefore, the embodiment not only can use other data sources to proofread the document and ensure the processing accuracy of the document, but also can use data calibration of other data sources to realize cross-verification of multi-source data of the office system and ensure the data consistency of multi-source data of the office system.

[0014] The data processing method based on multi-source data provided in the embodiment can be executed by a data processing device based on multi-source data. The data processing device based on multi-source data can be realized by software and / or hardware, and can be composed of two or more physical information or one physical information. For example, the data processing device based on multi-source data can be an office system or a background processing module of the office system. The background processing module can be deployed on the server side or the client side. Therefore, the data processing device based on multi-source data can also be a server side or a client side.

[0015] Further, the office system can be compatible with multiple functional systems, such as at least one of a document system, an email system, a knowledge base system, a space system, a meeting system, a contact list system, an instant messaging system, and a project system. Alternatively, in addition to being compatible with functional systems, the office system also has access permissions across multiple data sources, such as accessing some commonly used document platforms, email platforms, and meeting platforms to obtain data sources of these platforms.

[0016] The data processing device based on multi-source data is installed with at least one type of operating system, wherein the operating system includes but is not limited to an Android system, a Linux system and a Windows system. The data processing device based on multi-source data can install at least one application program based on the operating system, and the application program can be an application program provided by the operating system or an application program downloaded from a third-party device or server. In this embodiment, the data processing device based on multi-source data has at least one application program that can execute the data processing method based on multi-source data. The application program that executes the data processing method based on multi-source data can be an application program for opening and editing a document or a browser.

[0017] For the convenience of understanding, this embodiment takes an office system as an example to describe the subject that executes the data processing method based on multi-source data.

[0018] Figure 1 A flowchart of the data processing method based on multi-source data provided by this embodiment is given. Referring to FIG. 1, the data processing method based on multi-source data specifically includes the following steps. Figure 1 S110, obtaining target data of a first data source. S110, obtaining target data of a first data source.

[0019] The first data source is a content-editable data source to be processed, which can be any content-editable data source related to an office scenario, such as a document, a mail body, a mail attachment, a meeting minutes, personal space information, team space information, project information, calendar information and knowledge base data.

[0020] For example, when a user converts any content-editable data source in the office system into an editing mode or opens any content-editable data source, the office system identifies the data source in the editing mode or the opened data source as the first data source, and thus executes the data processing method of S110-S130 on the first data source.

[0021] Alternatively, the office system can obtain the first data source according to a preset processing task. The preset processing task can set the source range or identifier of the first data source, the processing requirement of the first data source, and the task execution time, etc. In addition to being manually created by the user on the office system, the processing task can also be automatically created in conjunction with other functional systems of the office system. When receiving a new message, completing a certain matter, or updating a locally managed data source, the other functional systems can create a corresponding processing task in the office system according to the new message, the completed matter, or the updated data source, so that the office system executes the processing task to perform data processing on the first data source associated with the message, matter, or data source, ensuring the consistency of the multi-source data of the office system. For example, when the online meeting system of the office system discusses the progress report of project A, the online meeting system can create a corresponding processing task in the office system for the progress report of project A, and the first data source of the processing task is the progress report of project A, and the processing requirement of the processing task is the meeting content. For another example, when the mail system of the office system receives the equipment procurement contract of project B, the mail body requires checking the data of the equipment procurement contract, and the mail system can create a corresponding processing task in the office system for the equipment procurement contract of project B, and the first data source of the processing task is the equipment procurement contract of project B, and the processing requirement is the mail body. For another example, when the knowledge base system of the office system modifies the version number of the service interface, the knowledge base system can create a corresponding processing task in the office system for the data source related to the service interface, and the first data source of the processing task is the data source related to the service interface, and the processing requirement of the processing task is to calibrate the version number of the service interface.

[0022] After obtaining the first data source, the office system obtains target data from the first data source. The target data is data to be proofread in the first data source. For example, when the user opens the first data source or converts the first data source to an editing mode using the office system, the office system displays the content page of the first data source on the front-end interface, and the user can manually select the target data of the first data source. The office system responds to the data selection operation and takes the selected data as the target data. For example, when the first data source is a contract document, the user wants to check the contract amount, delivery time, and contracting party in the contract document, and selects the contract amount, delivery time, and contracting party in the contract document and clicks the check control. The office system responds to the click operation triggered by the check control and takes the contract amount, delivery time, and contracting party as the target data.

[0023] In addition to being manually triggered, the target data can also be automatically obtained from the first data source by the office system. Specifically, the office system extracts a plurality of first key information as target data from the first data source. The first key information is information related to the core content in the first data source, which can be a keyword and a key sentence.

[0024] For example, when the first data source is a contract document, the core content of the contract document is legal provisions, amount, signatory party, time, etc., so these key information can be taken as the target data.

[0025] Alternatively, the distribution position of the core content of the first data source is determined according to the source type of the first data source, and the corresponding first key information is extracted from the first data source according to the distribution position of the core content. The source type includes email body, meeting minutes, personal space information, team space information, project information, calendar information and knowledge base data, etc. For example, the core content in the meeting minutes is the participants, the meeting title, the meeting highlights, etc. These core contents are distributed in the specific position of the meeting minutes according to the preset format, and the participants, the meeting title, the meeting highlights can be obtained from the specific position of the meeting minutes as the first key information.

[0026] In addition, if the office system obtains the corresponding first data source based on the preset processing task, the core content of the first data source can be determined according to the processing requirement of the processing task, and the first key information can be extracted from the first data source according to the core content. For example, when the first data source is a project plan report of a meeting discussion, the processing requirement is the meeting content, the core content of the project plan report is analyzed according to the meeting content, such as the online time of the project discussed in the meeting, the responsible person and the cooperation department, etc., and these contents are taken as the core content so as to extract the online time, the responsible person and the cooperation department, etc. from the first data source as the first key information.

[0027] However, the above-mentioned automatic acquisition methods of the first key information all need to pre-set the corresponding extraction rules, and the pre-operation is troublesome. Moreover, the extraction rules limit the use scene and the use object. Once the use scene and the use object are deviated, the adaptation degree of the extraction rules is low, resulting in that the extraction of the first key information is not comprehensive and accurate. In this regard, the entity words in the first data source can be extracted as the target data, and the entity words involve the name, the project name, the time and the version number, etc. The extraction of the entity words can be detected and extracted by using some open source entity detection model, which is suitable for various scenes and objects.

[0028] S120, obtaining the associated data of the target data from the second data source.

[0029] The second data source is a data source associated with the first data source, which can be a data source in a functional system used in an office scenario such as a document system, a mail system, a knowledge base system, a space system, a conference system, an address book system, an instant messaging system, and a project system, and can be a document, a mail, a knowledge base record, conference content, chat records, calendar information, project information, team space information, and personal space information. The second data source is not limited to whether it can be edited, that is, a data source that cannot be edited can also be a second data source.

[0030] The second data source can be selected by the user or automatically determined by the office system based on the first data source.

[0031] When the office system automatically determines the second data source, the office system can obtain the second data source associated with the first data source in each functional system according to the source information and / or the description information of the first data source. The source information refers to information related to the generation of the first data source, such as the functional system, project, space, author, and time to which the first data source belongs, and the description information refers to information that explains the content and purpose of the first data source, such as the name and note of the first data source. For example, when the first data source is a plan report of project A, all data sources related to project A, such as conference records, mail records, knowledge base records, project information, chat records, and space records, can be used as the second data source. For another example, the relevant data sources can be searched in each functional system of the office system based on the name and note of the first data source. For example, when the first data source is a plan report of project A, if an online conference A discussing the first data source is found in the conference system, the conference record of the online conference A is obtained as the second data source.

[0032] Optionally, in order to reduce the pressure of the office system querying the second data source and improve the processing efficiency of the first data source, the data space most closely associated with the first data source can be determined to obtain the second data source from the data space. Specifically, the second data source is determined from the preset space associated with the first data source. The preset space can be understood as a personal space, a team space, a mail space, a knowledge base space, and a meeting space, and the like, which are spaces for isolating other irrelevant data sources. In the case that the first data source has a corresponding belonging space, the belonging space of the first data source is taken as the associated preset space, and then the data source in the belonging space is taken as the second data source. In the case that the first data source does not have a corresponding belonging space, the information such as a project, an author, and a team associated with the first data source can be used to determine the associated space, so that the data source in the associated space is taken as the second data source. For example, when the first data source is the body of mail A, the first data source can be regarded as the data source of the mail A space, and the attachment of the mail A also belongs to the data source of the mail A space, so the attachment of the mail A can be taken as the second data source. For another example, the first data source is a report plan of a project A, and the project A is associated with a team space A, so the team space A can be determined as the associated space of the first data source, and then the data source in the team space A is taken as the second data source.

[0033] In the embodiment, the preset space associated with the first data source is determined to obtain the second data source highly related to the first data source from the preset space, so as to avoid blind acquisition of irrelevant data sources, improve the efficiency of associated data acquisition, reduce invalid data interference, and ensure the accuracy of subsequent data consistency detection.

[0034] Further, the more the associated data sources of the first data source, the more the interference items, and therefore, the similarity between the content semantics of the first data source and the content semantics of the associated data source can be matched to obtain the associated data source with high similarity as the second data source. The associated data source with the highest similarity can be taken as the second data source, or the associated data source with a similarity higher than a preset similarity threshold can be taken as the second data source.

[0035] After obtaining the second data source, the associated data of the target data can be obtained in the second data source. The associated data is a reference benchmark of the target data, and the associated data and the target data point to the same information dimension, such as the same project leader, the same project online time, and the same project service interface version number. It can be understood that in the office system, different data sources involve multiple collaborators, so that the data used to describe the same information dimension in different data sources is inconsistent, thereby affecting the collaboration of the office system. The embodiment aims to compare the target data with the associated data for consistency to determine the target data with an abnormality, so as to facilitate adjustment of the target data and / or the associated data, and ensure the consistency of the multi-source data of the office system.

[0036] When the associated data of the target data in the second data source is obtained, the target data in the second data source can be directly subjected to keyword matching or semantic matching, and the text with high keyword matching or high semantic matching can be taken as the associated data of the target data. However, the target data has few keywords and single semantic information, and directly using the target data to match the second data source can cause false matching, and the associated data with low relevance to the target data is matched, which affects the accuracy of subsequent consistency comparison. In this regard, the associated data in the second data source can be determined in combination with the target data and the context of the target data in the first data source.

[0037] Specifically, the context data of the target data in the first data source is obtained, and the associated data of the target data in the second data source is obtained based on the context data and the target data. For example, when the first data source is a document, the paragraph, chapter or container where the target data is located in the first data source can be taken as the context data of the target data, the context data of the associated data is matched in the second data source based on the context data of the target data, and the associated data is matched in the context data of the associated data based on the target data.

[0038] When the context data of the associated data is matched in the second data source, the semantic information of the context data of the target data can be obtained, the context data of the associated data is matched in the second data source according to the semantic information of the context data, or the corresponding keywords such as project name and project leader are matched in the second data source according to the keywords of the context data of the target data, and the paragraph, chapter or container corresponding to the keywords is taken as the context data of the associated data. When the associated data is matched in the context data of the associated data based on the target data, the keywords or semantic information of the target data can also be used to determine the associated data with keyword matching or semantic similarity in the context data of the associated data.

[0039] Alternatively, when the associated data of the target data in the second data source is obtained based on the context data and the target data, a prompt word is constructed based on the context data and the target data, the prompt word and the second data source are input into a preset large language model, and the preset large language model is guided by the prompt word to determine the associated data of the target data in the second data source and output. The prompt word can be "refer to the context description of text 1 in the file to find the text pointing to the same information as text 2 and output", text 1 points to the context data of the target data, text 2 points to the target data, and the file points to the second data source.

[0040] The embodiment introduces the context data of the target data to match the associated data in the second data source, improves the matching accuracy of the associated data, avoids obtaining redundant information irrelevant to the target data, improves the data processing efficiency while ensuring the data processing accuracy.

[0041] Optionally, if the target data is the first key information extracted from the first data source, a plurality of second key information can be extracted from the second data source, and the associated data of the target data is determined in the plurality of second key information. For example, the plurality of second key information can be extracted from the second data source in the same way as the first key information, and then the second key information with the same extraction rule as the first key information is taken as the associated data of the first key information. For example, the first key information is the contract amount, the delivery time and the contracting party in the first data source, and then the contract amount, the delivery time and the contracting party can be extracted from the second data source. Obviously, the contract amount in the first data source and the second data source is extracted based on the same extraction rule, which indicates that the extraction rule extracts data pointing to the same information dimension, and therefore the contract amount of the first data source and the contract amount of the second data source can be taken as a pair of target data and associated data. Similarly, the delivery time of the first data source and the delivery time of the second data source can be taken as a pair of target data and associated data, and the contracting party of the first data source and the contracting party of the second data source can be taken as a pair of target data and associated data.

[0042] The embodiment extracts the key information of the first data source and the key information of the second data source as the target data and the associated data respectively, focuses on the core content to be checked, avoids invalid processing of non-key information, and improves the data correction efficiency. In addition, the target data and the associated data pointing to the same information dimension are matched in the target data and the associated data, so as to ensure the accuracy of data matching.

[0043] It should be noted that the above method of matching the first key information and the second key information based on the extraction rule is limited to the case where the first key information and the second key information are obtained by the extraction rule. If the first key information and the second key information are entity words or key sentences extracted by the model, the extraction rule matching method is not suitable. In this regard, when the entity word or sentence extracted by the model is taken as the first key information or the second key information, a mapping relationship between the plurality of target data and the plurality of second key information can be established, and the second key information corresponding to the target data is determined as the associated data of the target data.

[0044] The mapping relationship refers to the relationship between the target data and the associated data pointing to the same information dimension. When the first key information and the second key information are sentences, the semantic similarity of the first key information and the second key information can be matched, and the first key information and the second key information with high semantic matching can be established to have a mapping relationship.

[0045] When the first key information and the second key information are entity words, an entity graph can be generated according to the first key information and the second key information, the first key information and the second key information are nodes of the entity graph, and an association relationship between the first key information and the second key information is an edge of the entity graph. The entity words include a document chapter, a meeting, a mail thread, a requirement item, a task, a version number, a person name, a department, a schedule event, and the like, and the edge includes a reference relationship, an update relationship, a decision relationship, an assignment relationship, an ownership relationship, an equivalence relationship, and the like. For example, the project name, the project leader name, the team name, the project online date, the project requirement, the project version number, and the like are extracted from the first data source and the second data source as the entity words. Based on the entity words, a decision relationship between the project and the project leader and an ownership relationship between the project leader and the team can be constructed. Then, the entity graph is constructed according to the entity words and the association relationship. From the entity node corresponding to the first key information, all the entity nodes corresponding to the second key information directly or indirectly associated with the first key information are traversed, and the edge attributes of each node are recorded. According to the association relationship type between the first key information and the corresponding second key information, whether the first key information and the second key information are in a mapping relationship is determined. When the association relationship between the first key information "Zhang San" and the first key information "project A" is a decision relationship, it indicates that Zhang San is the project leader of project A in the first data source. The association relationship between "project A" and the second key information "project A" is an equivalence relationship, and the association relationship between the second key information "Li Si" and the second key information "project A" is also a decision relationship, which indicates that Li Si is the project leader of project A in the second data source. Since Li San and Li Si are both project leaders of project A, they point to the same information dimension, and the two entity words can be used as target data and associated data in a mapping relationship. It can be seen that when the association relationship between the first key information and the directly associated second key information is an equivalence relationship, it is determined that the first key information and the directly associated second key information are in a mapping relationship. When the association relationship of each edge between the first key information and the indirectly associated second key information is symmetrical, it is determined that the first key information and the indirectly associated second key information are in a mapping relationship. In this embodiment, the entity graph is established by taking the first key information and the second key information as nodes, so that the second key information corresponding to the first key information is quickly searched in the entity graph, and the establishment efficiency of the mapping relationship is improved. Moreover, the field-level anomaly detection can be realized through the consistency detection of the entity words, and the detection accuracy is improved.

[0046] In this embodiment, the mapping relationship between the target data and the second key information is established, the associated data is quickly and accurately matched, the correspondence between the associated data and the target data is ensured, and a clear comparison basis is provided for subsequent consistency detection.

[0047] S130, consistency detection is performed on the association data and the target data to determine an abnormal problem of the target data.

[0048] For example, after obtaining the association data and the target data, consistency detection can be performed on the association data and the target data to detect whether the information of the association data and the target data is consistent, so as to determine the abnormal problem of the target data according to the detection result. It should be noted that the abnormal problem of the target data does not necessarily mean that the target data needs to be modified, but in the case that the information of the target data and the association data points to the same information dimension and is inconsistent, one of them has a problem, or the information inconsistency is caused by project requirements.

[0049] Optionally, when performing consistency detection on the association data and the target data, whether the information of the two is consistent can be directly compared. In the case that the information of the association data and the target data is consistent, it can be determined that the target data does not have an abnormal problem.

[0050] In the case that the information of the association data and the target data is inconsistent, it is determined that the abnormal problem of the target data is information conflict. For example, the target data is "person in charge Zhang San", and the association data is "person in charge Li Si". Obviously, the information dimension of the person in charge is inconsistent in the first data source and the second data source, indicating that the information of the person in charge of one of the data sources is in conflict, so that the target data of the person in charge Zhang San is marked as having an abnormal problem of information conflict with the association data.

[0051] Further, when the first data source is associated with multiple second data sources, the target data can determine an association data in multiple second data sources, so that the target data corresponds to multiple association data. At this time, the target data can be compared with each association data one by one to determine whether the target data and each association data are consistent. If the target data and each association data are consistent, it is determined that the target data does not have an abnormal problem. If the target data and any association data are inconsistent, it is determined that the target data has an abnormal problem of information conflict.

[0052] In the case that the association data lacks the target data, it is determined that the abnormal problem of the target data is information missing. For example, the target data is that the project will be online on December 30, and the association data is that the project will be online soon. Obviously, the information dimension of the project online date exists in the first data source, but does not exist in the second data source, indicating that the project online date of the second data source is missing, so that the target data of the project online on December 30 is marked as having an abnormal problem of information missing of the corresponding association data.

[0053] Based on the information conflict or information missing between the target data and the association data, the specific type of the abnormal problem is determined in the embodiment, so that the abnormal detection result is more clear and concrete, and the user can handle the abnormal problem specifically, and the data processing efficiency is improved.

[0054] It should be noted that in determining the associated data of the target data, there can be associated data that cannot be matched, which can also prove that the associated data of the target data is missing in the second data source. In this case, it can also be determined that the target data has an abnormal problem of information missing. For example, in constructing an entity graph, the first key information is not found in the second key information of the indirect association, and the second key information that satisfies the symmetry of the connection edge is not found, so it is determined that the first key information does not have corresponding mapping second key information, and further determines that the first key information has an abnormal problem of information missing.

[0055] After determining the abnormal problem of the target data, the target data or the associated data can be automatically modified based on the abnormal problem of the target data. Specifically, the target data is modified based on the associated data; or the associated data is modified based on the target data; or the target data is added in the associated data.

[0056] For example, when the target data has an abnormal problem of information conflict, the target data can be modified based on the associated data or the associated data can be modified based on the target data. Among them, the authority of the first data source and the second data source can be compared. When the authority of the first data source is higher than that of the second data source, the associated data is modified based on the target data. When the authority of the first data source is lower than that of the second data source, the target data is modified based on the associated data. The authority of the first data source and the second data source can be determined according to the timeliness and / or importance of the first data source and the second data source.

[0057] When the target data is modified based on the associated data, the associated data can replace the target data in the first data source. For example, the target data in the email body is "person in charge Zhang San", and the associated data in the email attachment is "person in charge Li Si". The "person in charge Zhang San" in the email body can be modified to "person in charge Li Si" to ensure that the information dimension of the person in charge in the modified email body and the email attachment is consistent. Similarly, when the associated data is modified based on the target data, the target data can replace the target data in the second data source.

[0058] If the target data corresponds to multiple associated data, and the multiple associated data are all different, the authority of the first data source and the second data source to which the multiple associated data belong can be compared. If the authority of the first data source is the highest, the associated data inconsistent with the target data can be modified based on the target data. If the authority of a second data source is the highest, the target data can be modified based on the associated data of the second data source, and even the associated data of other second data sources can be modified based on the associated data of the second data source.

[0059] When the target data has an abnormal problem of missing information, the target data can be added in the associated data. For example, when the target data in the project plan report is "project online on December 30", and the associated data in the project meeting minutes is "project is about to go online", the relevant information of "project online on December 30" can be added in the project meeting minutes to ensure that the information dimensions of the project plan report and the project meeting minutes pointing to the online date are consistent. Alternatively, when the target data does not have corresponding mapping associated data, the target data can be added to the second data source. For example, the target data in the project plan report is "project online on December 30", and there is no record of the project online time in the project calendar information. The information of "project online time is December 30" can be added in the project calendar information.

[0060] The embodiment supports three flexible modification modes to adapt to different abnormal problems, improves the modification universality, and further ensures the data consistency of various data sources in the office system, avoiding office collaboration failures caused by data conflicts.

[0061] In addition to automatically modifying the target data or the associated data, the abnormal problem of the target data can also be prompted to allow the user to manually modify the target data or the associated data under the prompt. For example, the office system can compile the abnormal problems of each target data of the first data source and the associated data into an abnormal table, and display the abnormal table of the first data source on a task result interface to prompt the user which target data of the first data source has an abnormal problem through the abnormal table. The task result interface is a result display interface of the corresponding processing task when the first data source is processed according to the preset processing task, and the user can view the abnormal problem of the target data of the corresponding first data source through the result display interface of each processing task.

[0062] When the content display interface of the first data source is in an open state, the abnormal problem of the target data can be directly associated and prompted in the content display interface of the first data source, so that the user can quickly modify the target data of the first data source based on the prompt. The content display interface is an interface for displaying the data content of the first data source. For example, the annotation information of the target data is generated according to the abnormal problem. The annotation information references the target data in the first data source and adds the text of the abnormal problem of the target data. For example, Figure 2 is one of the schematic diagrams of the content display interface of the first data source provided by the embodiment. As shown in Figure 2 , the content data of the first data source is displayed in the content area 11 of the content display interface, the target data is "Zhang San", and the annotation information of the target data is displayed in the annotation box 12. The annotation box displays that the abnormal problem of the target data is information conflict.

[0063] Or, according to the abnormal problem, generate revision information of the target data. Wherein, the revision information is the associated data displayed in the revision form behind or below the target data, and the associated data in the revision form is marked by underline or bracket. Figure 3 is a schematic diagram of a content display interface of a first data source provided by an embodiment of the present application. As shown in Figure 3 , the target data is "Zhang San", and the associated data is "Li Si". The abnormal problem of the target data is information conflict. The associated data corresponding to the conflict is directly displayed in the revision form behind the target data. The user can know that there is an abnormal problem of information conflict between the responsible person names of "Zhang San" and "Li Si".

[0064] Or, according to the abnormal problem, the target data is displayed differently in the first data source. For example, the target data with the abnormal problem of information conflict can be marked with a yellow highlight block, and the target data with the abnormal problem of information missing can be marked with a blue highlight block. Figure 4 is a schematic diagram of a content display interface of a first data source provided by an embodiment of the present application. As shown in Figure 4 , the target data is "Zhang San", and the associated data is "Li Si". The abnormal problem of the target data is information conflict. The associated data corresponding to the conflict is directly displayed in the revision form behind the target data. The user can know that there is an abnormal problem of information conflict between the responsible person names of "Zhang San" and "Li Si".

[0065] Or, the target data and the abnormal problem are added to the first comprehensive display area of the first data source. Wherein, the first comprehensive display area is an area in the content display interface of the first data source for displaying the abnormal problem of the first data source. The target data and the corresponding abnormal problem of the first data source are displayed in the first comprehensive display area, so that the user can know all the abnormal problems of the target data of the first data source at a glance. Figure 5 is a schematic diagram of a content display interface of a first data source provided by an embodiment of the present application. As shown in Figure 5 , the first comprehensive display area 13 is a sidebar of the content area 11. Two target data of the first data source have abnormal problems, which are "Zhang San" and "Project 12 / 20 online", respectively. The two target data and the corresponding abnormal problems can be summarized in the first comprehensive display area 13.

[0066] The office system can be compatible with the above-mentioned various prompt methods of prompting the abnormal problem in the content display interface. The user can select one of the methods according to the use habit. This embodiment provides diversified abnormal prompt methods to prompt the user the target data and the corresponding abnormal problem in the first data source when the first data source is in the content display state, adapts to the use habit of different users, and improves the user experience. Moreover, this embodiment visually displays the abnormal problem of the target data in the content display interface of the first data source, which is convenient for the user to quickly locate the abnormality and understand the abnormal situation, and improves the data processing efficiency.

[0067] Further, in the case of displaying the first comprehensive display area, the user can jump to the corresponding position of the target data in the first data source through the target data in the first comprehensive display area. Specifically, the user triggers a selection operation (such as a single click, double click, or long press) on the target data in the first comprehensive display area, and the office system jumps to the display page of the target data in the first data source in response to the selection operation on the target data in the first comprehensive display area. Referring to Figure 5 When the user clicks the target data "Zhang San" of the first comprehensive display area 13, the paragraph of "Zhang San" in the first data source can be obtained, and the content area is controlled to jump to the display page of "Zhang San" in the first data source according to the paragraph. At the same time, the first comprehensive display area is followed by display, and the user can modify or confirm that there is no abnormal problem in the target data in the first data source according to the abnormal problem in the first comprehensive display area. The embodiment can quickly locate the target data in the first data source through the target data in the first comprehensive display area, so as to quickly carry out the processing operation of the abnormal problem, ensure the coherence of the viewing and processing of the abnormal problem, and optimize the data processing process.

[0068] When the first data source has an associated preset space, the abnormal problem of the target data can also be prompted in the preset space associated with the first data source. For example, an abnormal prompt area associated with the first data source can be set in the space display interface of the preset space, and the abnormal problem of each target data of the first data source is displayed in the abnormal prompt area, so that the user can quickly confirm the abnormal problem of the target data of the first data source in the abnormal prompt area associated with the first data source. Figure 6 is one of the schematic diagrams of the space display interface of the team space provided by the embodiment of the present application. As Figure 6 shown, the abnormal prompt area 22 of the document 1 in the team space is displayed in association with the thumbnail 21 of the document 1, and the target data in the document 1 and the corresponding abnormal problem are displayed in the abnormal prompt area.

[0069] It should be noted that, except for some special spaces in which the first data source is in a content display state (such as the body of an email in an email space), most of the first data sources are in a content collection state and only display the corresponding thumbnail or file name in the space, and the abnormal prompt area associated with the first data source can help the user to confirm the abnormal problem of the first data source without opening the first data source. For the user, he can prioritize more urgent abnormalities in the abnormal problems of each data source in the space to ensure the timeliness of the office.

[0070] In addition to setting the abnormal problem in the abnormal prompt area corresponding to the first data source, the abnormal problems of all data sources in the preset space can be comprehensively displayed. Specifically, the abnormal problem is added to the second comprehensive display area of the preset space associated with the first data source, and the second comprehensive display area also includes abnormal problems of other data sources in the preset space. Figure 7 FIG. 2 is a schematic diagram of a space display interface of a team space provided by an embodiment of the present application. As shown in FIG. 2, the second comprehensive display area 23 simultaneously displays the target data and abnormal problems of the document 1, the document 2 and the document 3 in the team space, and the target data and abnormal problems of each document are distributed below the corresponding document. The distribution area is clear, and the user can determine the abnormal problem of each document in the second comprehensive display area 23. Figure 7

[0071] The embodiment of the present application sets the abnormal problem in the preset space associated with the first data source, so that the user can confirm whether the first data source has an abnormal problem without opening the first data source, avoiding the user opening the first data source without an abnormal problem to confirm the abnormal problem, and affecting the data source processing efficiency. Moreover, when the preset space is a team, a meeting or other multi-person collaboration space, the abnormal problem is prompted in the preset space to realize collaborative sharing of abnormal information, ensuring that all related collaborators know the abnormal problem, avoiding the abnormal problem being limited to a single user, and improving the efficiency of multi-person collaborative processing of the abnormal problem. In addition, the embodiment of the present application centrally displays the abnormal problems of multiple data sources in the preset space, realizes unified management and centralized viewing of abnormal information, and facilitates overall grasp of the abnormal situation of multiple data sources.

[0072] Further, the processing priority of each data source can be estimated based on the number or urgency of the abnormal problems of each data source in the preset space, and the sorting of each data source in the preset space and the sorting of each data source in the second comprehensive display area are adjusted according to the processing priority, so that the user can preferentially process the data source with higher priority.

[0073] ​On the basis of the above prompt abnormal problem, the abnormal problem and the abnormal evidence can be associated and prompted. The abnormal evidence is a second data source used to prove that the target data is inconsistent with the associated data. For example, the abnormal evidence of the target data can be generated according to the associated data and the second data source to which the associated data belongs, and the abnormal evidence and the abnormal problem are associated and prompted. The paragraph or chapter in which the associated data in the second data source is located can be taken as the abnormal evidence, and the associated data in the abnormal evidence is differentially marked. Alternatively, the associated data and the name of the second data source to which the associated data belongs can be directly taken as the abnormal evidence, and the name of the second data source in the abnormal evidence is set as an access interface of the second data source, so as to jump to the display interface of the second data source through the access interface. After the abnormal evidence is generated, the abnormal evidence and the abnormal problem are associated and prompted, so that the user can judge the reliability of the abnormal problem through the abnormal evidence and select the corresponding processing mode.

[0074] It should be noted that no matter how the abnormal problem is prompted, the abnormal evidence can be associated and prompted with the abnormal problem. For example, Figure 2 the abnormal problem and the abnormal evidence are displayed in the same batch note box 12, Figure 5 the abnormal problem and the abnormal evidence are displayed in the same cell of the first comprehensive display area 13, Figure 6 and Figure 7 the abnormal problem and the abnormal evidence are displayed below the corresponding target data. When the abnormal problem is displayed in the form of revision or difference, a floating icon 14 is arranged near the target data, and when the floating icon 14 is clicked, an abnormal prompt window of the target data can be expanded, and the abnormal problem and the abnormal evidence are displayed in the abnormal prompt window. For example, Figure 8 is the fifth schematic diagram of the content display interface of the first data source provided by the embodiment of the present application. As shown in Figure 8 when the user clicks the floating icon 14 of the target data, an abnormal prompt window 15 can be expanded at the original position of the floating icon 14, and the abnormal problem and the abnormal evidence are displayed in the abnormal prompt window 15.

[0075] The embodiment generates the corresponding abnormal evidence for the abnormal problem and associates and prompts the abnormal evidence, provides data support for the abnormal evidence, facilitates the user to check whether the abnormal problem is accurate through the abnormal evidence, and provides an objective basis for subsequent abnormal processing.

[0076] Further, the name of the second data source in the abnormal evidence points to an access interface of the second data source, and when the user triggers a selection operation on the name of the second data source in the abnormal evidence, a display page of the second data source corresponding to the abnormal evidence is opened in response to the selection operation on the second data source in the abnormal evidence. For example, the name of the second data source is saved in association with the access interface thereof, and when the selection operation of the second data source is triggered, the access interface saved in association with the name of the second data source is acquired, the original file of the second data source is acquired through the access interface, and the original file of the second data source is loaded through a pop-up window to display the content data of the second data source in the pop-up window. As shown in FIG. 8, the second data source is "XX file", which is marked by a down arrow, and the user can open the display page of the second data source by double-clicking "XX file". The present embodiment supports jumping from the abnormal evidence to the display page of the content data of the second data source, which facilitates the user to directly view the data source to which the associated data belongs and check the authenticity of the abnormal problem, thereby improving the efficiency of abnormal checking. Moreover, when the associated data needs to be modified, the display page of the content data of the second data source can be opened to facilitate the modification of the associated data of the second data source, thereby simplifying the process of modifying the associated data by the user. Figure 2

[0077] It should be noted that if the target data has multiple associated data, abnormal evidence of each associated data can be generated based on each associated data and the second data source to which the associated data belongs, so as to provide multiple evidences to facilitate the user to more accurately judge the authenticity of the abnormal problem and select a suitable processing mode.

[0078] Based on the above-mentioned abnormal problem, a modification suggestion of the abnormal problem can also be generated based on the abnormal problem of the target data and the associated data, and the abnormal problem and the modification suggestion are associated and prompted. The modification suggestion is a suggestion on how to modify the target data or the associated data.

[0079] For example, in the case of information conflict, a modification suggestion of modifying the target data based on the associated data is generated, or a modification suggestion of modifying the associated data based on the target data is generated. In the case where the target data corresponds to only one associated data, the authority of the first data source and the second data source can be compared. When the authority of the first data source is higher than that of the second data source, a modification suggestion of modifying the associated data based on the target data is generated, which indicates that the associated data in the second data source is replaced by the target data in the first data source, for example Figure 7 ​When document 2 is the first data source and document 3 is the second data source, the authority of document 2 is higher than that of document 3, the target data is "interface version V3", and the associated data is "interface version V1", so a modification suggestion of modifying "interface version V1 of document 3 to interface version V3" is generated. When the authority of the second data source is higher than that of the first data source, a modification suggestion of modifying the associated data based on the target data is generated, which indicates replacing the target data in the first data source with the associated data in the first data source, for example Figure 7 When document 1 is the first data source and document 3 is the second data source, the authority of document 1 is lower than that of document 3, the target data is "Zhang San", and the associated data is "Li Si", so a modification suggestion of modifying "Zhang San in document 1 to Li Si" is generated.

[0080] When the target data corresponds to multiple associated data, if the multiple associated data are consistent, when the authority of the first data source is higher than that of the second data source, a modification suggestion of modifying the associated data based on the target data is generated for each associated data, and when the authority of the first data source is lower than that of the second data source, a modification suggestion of modifying the target data based on the associated data is generated for the target data. If the multiple associated data are inconsistent, the first data source and the multiple second data sources are screened to obtain the second data source or the first data source with the highest authority, when the authority of the first data source is the highest, a modification suggestion of modifying the associated data based on the target data is generated for each associated data inconsistent with the target data. When the authority of the second data source is the highest, a modification suggestion of modifying the associated data based on the target data is generated for the target data, and a modification suggestion of modifying the associated data based on the associated data is generated for the associated data inconsistent with the associated data in the remaining second data sources. For example, the target data in contract 1 is "contract amount 30W", the associated data in contract 2 is "contract amount 40W", and the associated data in contract 3 is "contract amount 35W". If the authority of the first data source is the highest, a modification suggestion of modifying "contract amount 40W in contract 1 to 30W" can be generated for the associated data "contract amount 40W", and a modification suggestion of modifying "contract amount 35W in contract 3 to 30W" can be generated for the associated data "contract amount 35W". If the authority of contract 3 is the highest, a modification suggestion of modifying "contract amount 30W in contract 1 to 35W" can be generated for the target data "contract amount 30W", and a modification suggestion of modifying "contract amount 40W in contract 2 to 35W" can be generated for the associated data "contract amount 40W".

[0081] Alternatively, when the target data corresponds to multiple associated data and the multiple associated data are inconsistent, multiple candidate modification suggestions can be generated based on the multiple associated data, the candidate modification suggestions indicating modifying the target data to the corresponding associated data, and then the candidate modification suggestions are associated with the abnormal problem prompt of the target data, so as to facilitate the user to select a suitable modification suggestion from the multiple candidate modification suggestions.

[0082] When the abnormal problem is the missing of the associated information, a modification suggestion of adding the target data in the associated data is generated. The modification suggestion indicates adding the target data in the first data source to the associated data of the second data source. For example, Figure 7 When the Chinese document 1 is the first data source and the document 3 is the second data source, the abnormal problem of the target data "project on December 20" is that the associated data "the project is about to be put on the market" is missing the on-line time, and a modification suggestion of "adding the project on-line time of December 20 in the document 3" is generated.

[0083] The embodiment generates targeted modification suggestions for different abnormal types, improves the accuracy of the modification suggestions, ensures that the modification suggestions can solve the abnormal problems, improves the user adoption rate, and optimizes the user experience.

[0084] After the modification suggestion is generated, the modification suggestion and the abnormal problem can be associated and prompted. Referring to Figure 2 , the abnormal problem and the modification suggestion are displayed in the same batch note box 12. However, when the modification suggestion is based on the modification of the associated data based on the target data, the user lacks the support of the abnormal evidence corresponding to the associated data, and cannot confirm the rationality of the modification suggestion, so the modification suggestion, the abnormal evidence and the abnormal problem can be associated and prompted. Further, when the target data corresponds to multiple associated data, multiple abnormal evidences and multiple modification suggestions can be generated, the abnormal evidences corresponding to the multiple associated data can be summarized, the multiple modification suggestions can be summarized, and the summarized abnormal evidences and the summarized modification suggestions can be associated and prompted with the abnormal problem. Figure 9 is a third schematic view of a space display interface of a team space provided by an embodiment of the present application. As shown in Figure 9 , the target data in the contract 1 is "contract amount 30W", the associated data in the contract 2 is "contract amount 40W", the associated data in the contract 3 is "contract amount 35W", a modification suggestion of "modifying the contract amount 30W in the contract to 35W" is generated for the target data "contract amount 30W", and a modification suggestion of "modifying the contract amount 40W in the contract 2 to 35W" is generated for the associated data "contract amount 40W". The two modification suggestions can be summarized into one suggestion and displayed in the second comprehensive display area 23 in association with the abnormal problem, and the abnormal evidences corresponding to the contract 2 and the contract 3 can be summarized into one evidence and displayed in the second comprehensive display area 23 in association with the abnormal problem.

[0085] The embodiment generates a modification suggestion based on the abnormal problem and the associated data, so as to provide a clear abnormal processing direction for the user through the modification suggestion, and improve the abnormal processing efficiency. The abnormal problem is associated with the modification suggestion, so as to facilitate the user to quickly check the rationality of the modification suggestion, and further optimize the abnormal processing efficiency.

[0086] The modification suggestion can be a simple text content only for prompting, and the user can manually adjust the target data and / or the associated data according to the modification suggestion. In addition, the modification suggestion can also be a trigger interface of intelligent modification. When the user triggers an acceptance operation on the modification suggestion, the target data or the associated data can be modified based on the acceptance operation of the modification suggestion. Reference Figure 8 When the user double-clicks the modification suggestion, the acceptance operation on the modification suggestion can be triggered. In response to the acceptance operation of the modification suggestion, "Zhang San" is modified to "Li Si", and the corresponding revision trace can be retained after the modification for the user to check.

[0087] On the basis of the above-mentioned prompt of the abnormal problem, the abnormal problem and the credibility can be associated and prompted. The credibility represents the probability that the target data has the abnormal problem. Specifically, the credibility of the abnormal problem can be evaluated based on the association degree between the target data and the associated data, and / or the key degree of the associated data; and the abnormal problem and the credibility are associated and prompted. The association degree between the target data and the associated data can represent the probability that the target data and the associated data point to the same information dimension. The association degree can be determined based on the similarity of the semantic vectors between the target data and the associated data, or based on the edge length between the target data and the associated data in the entity graph. The longer the length is, the lower the association degree between the target data and the associated data is. When the association degree is higher, the probability that the target data and the associated data point to the same information dimension is higher, and the probability that the target data has the abnormal problem when the associated data is inconsistent with the target data is higher. The credibility of the abnormal problem can be directly evaluated based on the association degree between the target data and the associated data.

[0088] The key degree of the associated data represents the probability that the associated data is correct data in the office system. The key degree of the associated data can be determined based on the authority and timeliness of the second data source to which the associated data belongs and the importance of the associated data in the second data source. The higher the key degree of the associated data is, the higher the probability that the associated data is correct data is. The probability that the target data has the abnormal problem when the associated data is inconsistent with the target data is higher. The credibility of the abnormal problem can be directly evaluated based on the key degree of the associated data.

[0089] Of course, if the correlation between the target data and the related data is low, meaning the related data itself does not point to the same information dimension as the target data, then no matter how high the criticality of the related data is, it cannot prove whether the target data is correct. Similarly, if the criticality of the related data is very low, meaning the related data itself is incorrect, then no matter whether the target data and the related data point to the same information dimension, it cannot prove whether the target data is correct. Therefore, a weighted sum of the correlation between the target data and the related data, as well as the criticality of the related data, can be used to obtain the credibility of the anomaly.

[0090] It should be noted that when the target data corresponds to only one related data point, the anomaly can be identified based on the correlation between the target data and that related data point, and the criticality of that related data point. When the target data corresponds to multiple related data points, the credibility of each inconsistent related data point can be determined, and the highest, average, or median credibility value of each inconsistent related data point can be used as the credibility of the anomaly. Alternatively, the credibility of the anomaly can be determined based on the quantity, correlation, and criticality of each inconsistent related data point, where quantity is positively correlated with credibility.

[0091] After determining the credibility of the anomaly, you can associate the anomaly with its credibility level and provide a related suggestion. Figure 2 The anomaly and its credibility level are displayed in the same annotation box 12, allowing users to initially understand the authenticity of the anomaly based on the credibility level. Furthermore, if the credibility level is lower than a preset credibility threshold, it indicates that the target data does not have any anomalies, and the anomaly will not be highlighted. In other words, the anomaly of the target data will only be highlighted when the credibility level of the anomaly is higher than or equal to the preset credibility threshold.

[0092] This embodiment introduces a credibility assessment to provide users with a reference for judging the authenticity and importance of anomalies, reducing the ineffective handling of low-credibility anomalies. Furthermore, it can help users prioritize high-credibility anomalies, improving the efficiency of anomaly handling priority management.

[0093] The office system is related to enterprise management. In order to prevent important files in the enterprise from being modified by mistake or maliciously tampered with, some important files are set with corresponding modification permissions, so that only some specific personnel have the permission to modify the file content. For this embodiment, the first data source or the second data source can also be set with corresponding modification permissions, so that the office system can respond to the modification operation of the user on the first data source or the second data source when the user has the modification permission on the first data source or the second data source. Specifically, the office system receives the modification operation on the target data or the associated data, and according to the permission information of the first data source or the second data source, the modification operation is checked for permission; in the case where the permission check is passed, the target data or the associated data is modified according to the modification operation. The modification operation can be an operation of manually modifying the target data or the associated data by the user, or an operation of automatically modifying the target data or the associated data, or an operation of accepting the modification suggestion by the user. The permission information can include personal account, space identifier or project identifier, etc. that have the modification permission of the corresponding data source. The space identifier is the identifier of the space to which the data source belongs, so as to confirm that the user has the modification permission of the data source when the user triggering the modification operation comes from the user of the space to which the data source belongs. The project identifier is the identifier of the project to which the data source corresponds, and the user triggering the modification operation comes from the project group of the project to which the data source belongs, so as to confirm that the user has the modification permission of the data source. When the user triggers the modification of the target data, the modification operation of the target data is generated based on the personal account, the space identifier of the space to which the user belongs, or the project identifier of the project group to which the user belongs. The personal account, the space identifier or the project identifier, etc. are obtained from the modification operation, and it is determined whether the permission information of the first data source includes the personal account, the space identifier or the project identifier, etc. If it includes, it is determined that the permission check is passed, and if it does not include, it is determined that the permission check is not passed. In the case where the permission check is passed, the target data of the first data source is modified to the associated data. In the case where the permission check is not passed, the user is prompted that he does not have the operation permission to modify the first data source. The modification permission verification process of the second data source is the same, and is not described here. This embodiment adds a permission check mechanism to check the permission of the modification operation, prevents users without permission from modifying the first data source or the second data source, ensures the security of the first data source or the second data source, avoids the data from being modified by mistake or maliciously, and meets the demand of the office scene for data security management.

[0094] When the first data source or the second data source belongs to a multi-person collaborative data source, it is necessary to ensure that all collaborators are timely aware of data change, so as to avoid subsequent work failure caused by information asymmetry. To this end, after the first data source or the second data source is modified, a corresponding modification notification can be sent to all collaborators of the first data source or the second data source. Specifically, the collaborators of the first data source can be determined according to the source information of the first data source, and the modification notification of the first data source is sent to the collaborators; or the collaborators of the second data source are determined according to the source information of the second data source, and the modification notification of the second data source is sent to the collaborators. The modification notification is a notification for informing the collaborators of the corresponding data source of the modification of the data source, which can include one modified content of the data source, or all modified contents of the data source. The source information of the data source can indicate which project, which space or which department the data source comes from, so that the relevant personnel of the corresponding project, space or department can be regarded as the collaborators of the data source. When the office system modifies the first data source, the abnormal problem, the abnormal evidence and the modified target data of the target data in the first data source are generated to generate a modification notification, and the modification notification is sent to the relevant personnel of the project, space or department to which the first data source belongs, so that the relevant personnel can know the modification of the first data source in the first time. The modification notification process of the second data source is the same, and will not be described here. The modification notification mechanism is added in this embodiment to push the modification notification to the collaborators after the modification of the data source, so as to guarantee the consistency of the collaborative modification of the multi-source data and improve the team collaboration efficiency.

[0095] In summary, the data processing method based on multi-source data provided by the embodiment of the application is to obtain the target data of the first data source as the to-be-audited content, obtain the associated data of the target data in the second data source to introduce the calibration support of the multi-source data, compare the consistency of the associated data and the target data to accurately locate the abnormal problem of the target data, realize the cross verification of the multi-source data, solve the problem of lack of proofreading of office data in the prior art, improve the comprehensiveness and accuracy of abnormal identification from the data source level, provide a reliable basis for subsequent abnormal processing, improve the accuracy of data processing, and be suitable for multi-data source collaborative scenarios, realize effective verification of key information in the collaborative scenario, and guarantee the data consistency of the multi-data source.

[0096] On the basis of the above-mentioned embodiments, Figure 10 The structure diagram of the data processing device based on multi-source data provided by the embodiment of the application is shown in FIG. 1. Referring to FIG. 1, Figure 10 The data processing device based on multi-source data provided by the embodiment of the application specifically includes: a target data acquisition module 31, an associated data acquisition module 32, and an abnormal problem determination module 33.

[0097] The target data acquisition module 31 is configured to acquire target data of a first data source. The association data acquisition module 32 is configured to acquire the association data of the target data from the second data source; The abnormal problem determination module 33 is configured to perform consistency detection on the association data and the target data, and determine the abnormal problem of the target data.

[0098] On the basis of the above-mentioned embodiments, the data processing apparatus based on multi-source data further comprises an abnormal modification module configured to modify the target data based on the association data after determining the abnormal problem of the target data; or modify the association data based on the target data; or add the target data in the association data.

[0099] On the basis of the above-mentioned embodiments, the association data acquisition module 32 comprises a second data source determination sub-module configured to determine the second data source from the preset space associated with the first data source before acquiring the association data of the target data from the second data source.

[0100] On the basis of the above-mentioned embodiments, the association data acquisition module 32 comprises a context acquisition sub-module configured to acquire the context data of the target data in the first data source; and a first association data acquisition sub-module configured to acquire the association data of the target data in the second data source based on the context data and the target data.

[0101] On the basis of the above-mentioned embodiments, the target data acquisition module 31 comprises a target data acquisition sub-module configured to extract a plurality of first key information as the target data in the first data source; and correspondingly, the association data acquisition module 32 comprises a second association data acquisition sub-module configured to extract a plurality of second key information in the second data source, and determine the association data of the target data in the plurality of second key information.

[0102] On the basis of the above-mentioned embodiments, the second association data acquisition sub-module comprises a mapping relationship establishment unit configured to establish a mapping relationship between the plurality of target data and the plurality of second key information; and an association data acquisition unit configured to determine the second key information corresponding to the target data as the association data of the target data.

[0103] On the basis of the above-mentioned embodiments, the abnormal problem determination module 33 comprises a conflict problem determination sub-module configured to determine that the abnormal problem of the target data is information conflict in the case that the information of the association data and the target data is inconsistent; and / or a missing problem determination sub-module configured to determine that the abnormal problem of the target data is missing association information in the case that the association data lacks the target data.

[0104] On the basis of the above-mentioned embodiments, the data processing apparatus based on multi-source data further comprises: a first abnormality prompt module configured to, after determining the abnormality problem of the target data, generate annotation information of the target data according to the abnormality problem; or generate revision information of the target data according to the abnormality problem; or differentially display the target data in the first data source according to the abnormality problem; or add the target data and the abnormality problem to a first comprehensive display area of the first data source.

[0105] On the basis of the above-mentioned embodiments, the data processing apparatus based on multi-source data further comprises: a target data display module configured to, after adding the target data and the abnormality problem to the first comprehensive display area of the first data source, in response to a selection operation on the target data in the first comprehensive display area, jump to a display page of the target data in the first data source.

[0106] On the basis of the above-mentioned embodiments, the data processing apparatus based on multi-source data further comprises: a second abnormality prompt module configured to, after determining the abnormality problem of the target data, generate abnormality evidence of the target data according to the associated data and the second data source to which the associated data belongs; and associate and prompt the abnormality evidence and the abnormality problem.

[0107] On the basis of the above-mentioned embodiments, the data processing apparatus based on multi-source data further comprises: a second data source display module configured to, after associating and prompting the abnormality evidence and the abnormality problem, in response to a selection operation on the second data source in the abnormality evidence, open a display page of the second data source corresponding to the abnormality evidence.

[0108] On the basis of the above-mentioned embodiments, the data processing apparatus based on multi-source data further comprises: a third abnormality prompt module configured to, after determining the abnormality problem of the target data, prompt the abnormality problem of the target data in a preset space associated with the first data source.

[0109] On the basis of the above-mentioned embodiments, the third abnormality prompt module comprises: a comprehensive prompt sub-module configured to add the abnormality problem to a second comprehensive display area of the preset space associated with the first data source, and the second comprehensive display area further comprises abnormality problems of other data sources in the preset space.

[0110] On the basis of the above-mentioned embodiments, the data processing apparatus based on multi-source data further comprises: a fourth abnormality prompt module configured to, after determining the abnormality problem of the target data, generate a modification suggestion for the abnormality problem based on the abnormality problem of the target data and the associated data; and associate and prompt the abnormality problem and the modification suggestion.

[0111] On the basis of the above-mentioned embodiments, the fourth abnormality prompting module comprises: a first suggestion generation submodule configured to, in the case of an information conflict as the abnormality problem, generate a modification suggestion for modifying the target data based on the associated data, or generate a modification suggestion for modifying the associated data based on the target data; and / or a second suggestion generation submodule configured to, in the case of missing associated information as the abnormality problem, generate a modification suggestion for adding the target data in the associated data.

[0112] On the basis of the above-mentioned embodiments, the data processing apparatus based on multi-source data further comprises: a fifth abnormality prompting module configured to, after determining the abnormality problem of the target data through consistency detection on the associated data and the target data, evaluate the credibility of the abnormality problem based on the correlation degree between the target data and the associated data, and / or the key degree of the associated data; and associate and prompt the abnormality problem and the credibility.

[0113] On the basis of the above-mentioned embodiments, the data processing apparatus based on multi-source data further comprises: a modification permission verification module configured to, after determining the abnormality problem of the target data, in the case of receiving a modification operation on the target data or the associated data, perform permission verification on the modification operation according to the permission information of the first data source or the second data source; and in the case of passing the permission verification, modify the target data or the associated data according to the modification operation.

[0114] On the basis of the above-mentioned embodiments, the data processing apparatus based on multi-source data further comprises: a modification notification module configured to, after modifying the target data or the associated data according to the modification operation, determine the collaborators of the first data source according to the source information of the first data source, and send a modification notification of the first data source to the collaborators; or determine the collaborators of the second data source according to the source information of the second data source, and send a modification notification of the second data source to the collaborators.

[0115] The data processing apparatus based on multi-source data provided by the embodiments of the present application can be used to execute the data processing method based on multi-source data provided by the above-mentioned embodiments, and has the corresponding functions and beneficial effects.

[0116] The data processing apparatus based on multi-source data provided by the embodiments of the present application can be used to execute the data processing method based on multi-source data provided by the above-mentioned embodiments, and has the corresponding functions and beneficial effects.

[0117] Figure 11 is a structural schematic diagram of a data processing device based on multi-source data provided by an embodiment of the present application, referring to Figure 11 The data processing device based on multi-source data comprises a processor 41, a memory 42, a communication device 43, an input device 44 and an output device 45. The number of the processor 41 in the data processing device based on multi-source data can be one or more, and the number of the memory 42 in the data processing device based on multi-source data can be one or more. The processor 41, the memory 42, the communication device 43, the input device 44 and the output device 45 of the data processing device based on multi-source data can be connected through a bus or other means.

[0118] The memory 42 as a computer readable storage medium can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the data processing method based on multi-source data of any embodiment of the present application (for example, the target data acquisition module 31, the associated data acquisition module 32 and the abnormal problem determination module 33 in the data processing device based on multi-source data). The memory 42 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system and application programs required by at least one function; the data storage area can store data created according to the use of the device, etc. In addition, the memory 42 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device or other non-volatile solid-state storage device. In some examples, the memory can further include a memory remotely arranged with respect to the processor, which can be connected to the device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.

[0119] The communication device 43 is used for data transmission.

[0120] The processor 41 executes various functional applications and data processing of the device by running the software programs, instructions and modules stored in the memory 42, that is, implements the above-mentioned data processing method based on multi-source data.

[0121] The input device 44 can be used to receive input digital or character information, and generate key signal input related to user settings and function control of the device. The output device 45 can include a display device such as a display screen.

[0122] The data processing device based on multi-source data provided above can be used to execute the data processing method based on multi-source data provided by the above-mentioned embodiments, and has corresponding functions and beneficial effects.

[0123] The embodiment of the present application further provides a storage medium comprising computer executable instructions, which, when executed by a computer processor, are used to perform a data processing method based on multi-source data, the data processing method based on multi-source data comprising: obtaining target data of a first data source; obtaining associated data of the target data from a second data source; performing consistency detection on the associated data and the target data to determine an abnormal problem of the target data.

[0124] Storage medium - any of various types of memory devices or storage devices. The term "storage medium" is intended to include an installation medium, e.g., a CD-ROM, floppy disks, or tape device; a computer system memory or random access memory such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; or a non-volatile memory such as a magnetic medium (e.g., a hard drive or optical storage); registers or other similar types of memory elements, etc. The memory medium can also include other types of storage medium and combinations thereof. In addition, the memory medium can reside in a first computer system's main memory, or in a second different computer system's memory, and the second computer system can provide the data to the first computer system over a network (e.g., the Internet) for execution by the first computer system. The term "memory medium" should be taken to include a single medium or multiple media that store the at least one set of instructions. As used herein, a "processor" includes any hardware system, hardware

[0125] Of course, the storage medium provided by the embodiment of the present application comprises computer executable instructions, which are not limited to the data processing method based on multi-source data as above, but can also perform the related operations in the data processing method based on multi-source data provided by any embodiment of the present application.

[0126] The data processing device based on multi-source data, the storage medium and the data processing equipment based on multi-source data provided in the above embodiments can perform the data processing method based on multi-source data provided by any embodiment of the present application, and the technical details not described in detail in the above embodiments can be referred to the data processing method based on multi-source data provided by any embodiment of the present application.

[0127] The above merely describes the preferred embodiments of the present application and the technical principles applied. The present application is not limited to the specific embodiments herein, and various obvious changes, modifications and replacements made by those skilled in the art without departing from the scope of the present application shall not be excluded. Therefore, although the present application is described in more detail through the above embodiments, the present application is not limited to the above embodiments, and more other equivalent embodiments can be included without departing from the concept of the present application, and the scope of the present application is determined by the scope of the claims.

Claims

1. A data processing method based on multi-source data, characterized in that, include: Obtain the target data from the first data source; Obtain the associated data of the target data from the second data source; A consistency check is performed between the associated data and the target data to identify any anomalies in the target data.

2. The method according to claim 1, characterized in that, After determining the anomaly in the target data, the following steps are included: Modify the target data based on the associated data; or... Modify the associated data based on the target data; or... Add the target data to the associated data.

3. The method according to claim 1, characterized in that, Before obtaining the associated data of the target data from the second data source, the method further includes: The second data source is determined from the preset space associated with the first data source.

4. The method according to claim 1, characterized in that, The step of obtaining the associated data of the target data from the second data source includes: Obtain the context data of the target data from the first data source; Based on the context data and the target data, obtain the associated data of the target data from the second data source.

5. The method according to claim 1, characterized in that, The target data obtained from the first data source includes: Extract multiple key pieces of information from the first data source as target data; Accordingly, obtaining the associated data of the target data from the second data source includes: Extract multiple second key information from the second data source, and determine the associated data of the target data from the multiple second key information.

6. The method according to claim 5, characterized in that, The step of extracting multiple second key information from the second data source and determining the associated data of the target data from the multiple second key information includes: Establish a mapping relationship between multiple target data and multiple pieces of the second key information; The second key information corresponding to the target data is determined as the associated data of the target data.

7. The method according to claim 1, characterized in that, The step of performing consistency detection between the associated data and the target data to determine anomalies in the target data includes: In cases where the related data and the target data are inconsistent, the anomaly in the target data is determined to be an information conflict; and / or, If the target data is missing from the associated data, the anomaly of the target data is determined to be missing associated information.

8. The method according to claim 1, characterized in that, After identifying the anomaly in the target data, the following is also included: Generate annotation information for the target data based on the aforementioned anomaly; or... Generate revision information for the target data based on the aforementioned anomaly; or... Based on the aforementioned anomaly, the target data is displayed differentially in the first data source; or... Add the target data and the anomaly to the first comprehensive display area of ​​the first data source.

9. The method according to claim 8, characterized in that, After adding the target data and the anomaly to the first comprehensive display area of ​​the first data source, the method further includes: In response to the selection operation of target data in the first comprehensive display area, the user is redirected to the display page of the target data in the first data source.

10. The method according to claim 1, characterized in that, After identifying the anomaly in the target data, the following is also included: Based on the associated data and the second data source to which the associated data belongs, generate evidence of anomalies in the target data; The abnormal evidence and the abnormal problem will be associated and highlighted.

11. The method according to claim 10, characterized in that, After associating the abnormal evidence with the abnormal issue, the method further includes: In response to the selection of a second data source in the abnormal evidence, the display page of the second data source corresponding to the abnormal evidence is opened.

12. The method according to claim 1, characterized in that, After identifying the anomaly in the target data, the following is also included: In the preset space associated with the first data source, an error message is displayed regarding the target data.

13. The method according to claim 12, characterized in that, The step of alerting the target data for abnormal issues in the preset space associated with the first data source includes: The aforementioned abnormal issues are added to the second comprehensive display area of ​​the preset space associated with the first data source. The second comprehensive display area also includes abnormal issues from other data sources in the preset space.

14. The method according to claim 1, characterized in that, After identifying the anomaly in the target data, the following is also included: Based on the anomalies in the target data and related data, generate modification suggestions for the anomalies. The abnormal issue and the suggested modification will be linked together and displayed.

15. The method according to claim 14, characterized in that, The process of generating modification suggestions for the anomalies based on the target data and related data includes: In the case where the anomaly is an information conflict, a modification suggestion is generated based on the associated data to modify the target data, or, a modification suggestion is generated based on the target data to modify the associated data; and / or, If the anomaly is due to missing related information, a modification suggestion is generated to add the target data to the related data.

16. The method according to claim 1, characterized in that, After performing consistency checks on the associated data and the target data to determine any anomalies in the target data, the method further includes: The credibility of the anomaly is assessed based on the correlation between the target data and the associated data, and / or the criticality of the associated data; The abnormal issue and the credibility level will be linked together to provide a notification.

17. The method according to claim 1, characterized in that, After identifying the anomaly in the target data, the following is also included: Upon receiving a modification operation on the target data or the associated data, the modification operation is subject to permission verification based on the permission information of the first data source or the second data source. If the permission verification passes, modify the target data or the associated data according to the modification operation.

18. The method according to claim 17, characterized in that, After modifying the target data or the associated data according to the modification operation, the method further includes: Based on the source information of the first data source, determine the collaborators of the first data source, and send a modification notification of the first data source to the collaborators; or, Based on the source information of the second data source, identify the collaborators of the second data source and send a modification notification of the second data source to the collaborators.

19. A data processing device based on multi-source data, characterized in that, include: The target data acquisition module is configured to acquire target data from the first data source. The associated data acquisition module is configured to acquire associated data of the target data from a second data source; The anomaly detection module is configured to perform consistency checks on the associated data and the target data to determine anomalies in the target data.

20. A data processing device based on multi-source data, characterized in that, include: One or more processors; A memory that stores one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the data processing method based on multi-source data as described in any one of claims 1-18.

21. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the data processing method based on multi-source data as described in any one of claims 1-18.