Task information extraction system and method based on labels and medium

By introducing a tag-based task information extraction system into the project management tool, the logical association and two-way synchronization between tasks and documents are achieved, and the problems of insufficient depth of linkage between tasks and documents are solved and the complex operation are improved, and the collaboration efficiency and intuitiveness of project management are improved.

CN120104773APending Publication Date: 2025-06-06SUZHOU BLACK CORAL SOFTWARE CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510172893.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In existing project management tools, the linkage between tasks and documents is insufficient, the synchronization efficiency is low, and the contextual relationship between tasks in the document is lacking, resulting in reduced collaboration efficiency and complex operation.

Method used

It provides a tag-based task information extraction system, which realizes the logical association between tasks and documents through the tag management module, document management module and embed module, and supports two-way synchronization. Users can jump to the tag management module with one click to view detailed information through the task tag embedded in the document, and directly modify task attributes through the tag management module.

Benefits of technology

It realizes the dynamic integration of tasks and documents, improves collaboration efficiency, reduces the time for task status review, information synchronization and content updates, and makes project management more intuitive and efficient, meeting the actual needs of complex project collaboration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104773A_ABST
    Figure CN120104773A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of project management, in particular to a tag-based task information extraction system and method and a medium. The tag-based task information extraction system comprises a tag management module, a task information extraction module and a task information extraction module, wherein the tag management module is used for creating and managing tags; the document management module is used for creating a document and editing and storing the created document; the embedding module is used for connecting the document management module with the label management module, so that the document information is dynamically bound with the task attributes; and the bidirectional synchronization module is used for ensuring real-time updating of the document information and the task attributes. Through the technical scheme, the problems that the linkage depth of the task and the document is insufficient, the synchronization efficiency is low, the context association of the task in the document is lacked and the operation is complex are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of project management, and in particular to a system, method and medium for extracting task information based on tags. Background Art

[0002] In traditional project management tools, task management and document management are usually two independent modules, lacking an effective association mechanism. Users need to manually switch interfaces to find relevant information, which can easily cause data inconsistency or omission of key information. Even if tasks and documents are associated in some way (such as hyperlinks), the document content will not be updated synchronously when the task status changes. Users need to manually adjust the document content, which increases the complexity of operations and easily leads to information lag. In existing systems, tasks are usually only associated with documents through hyperlinks or notes, and the context of the task in the document cannot be displayed, making it difficult for task executors to accurately understand the background and requirements of the task; users switch frequently between task management and document management, and the operation steps are complicated, which reduces work efficiency.

[0003] Existing task management tools: such as Jira, Asana, Trello, etc., provide task creation, assignment, status update and team collaboration functions;

[0004] Document management tools: such as Confluence, Google Docs, Notion, etc., which support document creation, editing, storage and sharing;

[0005] Integrated tools: such as Notion and ClickUp, attempt to integrate task management and document management to a certain extent, but the depth of linkage between the two is limited, and they are usually linked by hyperlinks.

[0006] (1) Technical features and shortcomings of Jira and Confluence integration:

[0007] Technical features: Jira and Confluence are integrated through plug-ins, which allow users to embed Jira issue cards in Confluence documents. Users can view basic information about the task and jump to Jira to view task details. When the task status is updated, the card information will be refreshed.

[0008] Technical deficiencies: In the current integration between Jira and Confluence, users cannot directly convert text descriptions in Confluence into trackable Jira task cards. Specifically, users need to manually copy the to-do content in Confluence to the Jira system, and then return to the Confluence page to insert the task cards that already exist in Jira. This process is cumbersome and error-prone, and lacks efficient automation support. In addition, the way documents and tasks are associated is relatively simple, lacking in deep linkage and flexibility, resulting in reduced collaboration efficiency.

[0009] (2) Technical features and shortcomings of Notion:

[0010] Technical features: Notion supports users to embed Notion database views in documents for task management. Notion databases can be viewed as a collection of tasks of the same type or a complete collection of all tasks under the same project. This feature allows users to easily manage and track various tasks in documents. Notion does not support inserting an entry in a database into a document.

[0011] Technical deficiencies: When a document (such as meeting minutes) contains multiple types of to-do items (for example, preliminary proposals, meeting arrangements, specific tasks, risk items, etc.) and each type of to-do item has different attributes and management processes, Notion's current solution has limitations. Users need to manually copy the link of the to-do item and add it to the Notion database to generate the corresponding entry, and then insert these entries into the to-do section of the document. This process is cumbersome and inefficient, and lacks automated processing and integrated management of different types of to-do items.

[0012] (3) Technical features and shortcomings of Google Docs:

[0013] Technical features: Google Docs supports adding task annotations to a paragraph of text in a document and displays the task status through integration with Google Tasks or Google Keep. However, task data updates need to be completed in an independent task management tool.

[0014] Technical deficiencies: The synchronization between task annotations and document content is one-way, and task updates are not automatically reflected in the document; there is a lack of task context information display and interactive functions, and jump operations are complicated.

[0015] (4) Technical features and shortcomings of ClickUp:

[0016] Technical features: ClickUp provides an integrated task management and document management solution, supporting embedding task lists in documents, and automatically displaying the document when the task status is updated. However, the embedded task cannot jump back to the original task page. At the same time, if the text in the document is used to create a task and perform an undo operation, the integrity of the content will be affected when the related task is deleted.

[0017] Technical deficiencies: The task functions in the document are relatively simple. In addition to displaying the task title and status, the project name is not displayed, and there is a lack of contextual information and deeper interactive functions. This leads to a relatively limited interactive experience of tasks in the document, and it is impossible to effectively display the task background or deeply manage the multi-dimensional needs of tasks. Summary of the invention

[0018] In order to solve the problems in the related art, the present invention provides a tag-based task information extraction system, method and medium, which solves the problems of insufficient linkage depth between tasks and documents, low synchronization efficiency, lack of contextual association of tasks in documents and complex operations.

[0019] In order to solve the above problems, the following technical solutions are provided:

[0020] A tag-based task information extraction system of the present invention includes:

[0021] The tag management module is used to create tags, and edit, assign, update the status of and store the created tags. Users can add tags to document information to convert it into a manageable object, and automatically generate related fields and task attributes to achieve structured management of tags;

[0022] A document management module, which is used to create documents, edit and store the created documents;

[0023] An embedding module is used to connect the document management module with the tag management module, so that the document information and the task attributes are dynamically bound. The tags embedded in the document can directly reference the relevant fields and task attributes in the tag management module to achieve the logical association between the task and the document. The user can jump to the tag management module to view the detailed information with one click through the task tags embedded in the document, thus achieving the mutual jump function. The embedding module supports adding tags and withdrawing tags through the document management module page;

[0024] Bidirectional synchronization module: When a user updates a task attribute, the operation triggers the real-time synchronization mechanism of the bidirectional synchronization module. The bidirectional synchronization module sends the update request to the back-end server through the communication protocol. The back-end server writes the changed data into the cache and pushes it to all participating user sessions through the real-time synchronization mechanism, ensuring that all users can view the changed content in real time and keep the interface consistent. At the same time, when other users update, the changes will be synchronized to the current user interface, realizing bidirectional synchronization and ensuring that document information and task attributes are updated in real time between multiple users and devices.

[0025] Through the above scheme, the association between the document management module and the tag management module is realized through the embedded module, which solves the problem of disconnection between tasks and documents, enables tasks to be embedded in documents and keep dynamically updated, thereby realizing the connection between tasks and documents in project management, thereby solving the problem of lack of context association of tasks in documents; users can jump to the tag management module to view detailed information with one click through the task tag embedded in the document, and at the same time allow the task attributes to be directly modified through the tag management module, which is conducive to simplifying the operation process between tasks and documents; through the embedded design of task tags and strict identification binding, the uniqueness and relevance of task attributes are ensured, while avoiding errors caused by manual input and improving the reliability of collaboration; through the two-way synchronization module, task attributes and document information are synchronized in real time to ensure the consistency and timeliness of information, effectively improve the efficiency of collaboration, and through the dynamic integration of tasks and documents, the time for task status consultation, information synchronization and content update is significantly reduced, making project management more intuitive and efficient, meeting the actual needs of complex project collaboration, thereby solving the problem of complex operation.

[0026] By identifying the task tags embedded in the document, tasks in the tag management module are generated. The task is traced back to the document ID and name, and the document page can be jumped to.

[0027] The task attributes include title, label, and status.

[0028] In the mutual jump function, it is allowed to directly modify the task attributes through the tag management module.

[0029] The bidirectional synchronization module synchronizes task attributes and document information in real time using a unique identifier.

[0030] The algorithm process for implementing the logical association between tasks and documents in the embedded module is:

[0031] The text content in the document and the task-related fields in the tag management module are represented by TF-IDF vectors. After the text content in the document is cleaned, segmented, and standardized, the TF-IDF value of each word is calculated.

[0032] TF represents the frequency of a word t appearing in a document d.

[0033]

[0034] IDF represents the importance of a word t in the entire document set D.

[0035]

[0036] TF-IDF is the product of TF and IDF.

[0037] TF-IDF(t,d,D)=TF(t,d)×IDF(t,D);

[0038] Where t represents a word, d represents a document, and D represents the entire document set;

[0039] Aggregate the TF-IDF values ​​of all words into a vector v,

[0040] v=[TF-IDF(t1,d,D),TF-IDF(t2,d,D)...];

[0041] Calculate the L2 norm of the vector v to obtain a high-dimensional vector Norm(d),

[0042]

[0043] Among them, ∥v∥ 2 is the L2 norm of the vector;

[0044] The text content of the final document is represented as a high-dimensional vector Norm(d).

[0045] Use cosine similarity to measure the similarity between documents and tags, calculate the cosine similarity between documents and all tags, select the most relevant tag as the task tag, and logically associate and bind the document with the tag. The formula is as follows:

[0046]

[0047] Where d is the TF-IDF vector of the document and l is the TF-IDF vector of the tag;

[0048] If no tag meets the association conditions, a new tag is created based on the text content of the document or its summary, and re-associated and bound.

[0049] The bidirectional synchronization module implements the synchronization mechanism in the following way:

[0050] S(D,N,F,C)=αD+βN+γF+δC+∈;

[0051] Among them, α, β, γ, δ are weight factors, indicating the influence of each parameter on the overall performance, and ∈ represents the contribution of other uncontrollable factors;

[0052] The objective function S(D,N,F,C) represents the overall performance of the system.

[0053] Where D∈[0,1] represents the proportion of data update, N is the number of clients (a positive integer), F∈[0,1] represents the update frequency ratio, and C represents the choice of cache strategy. Different values ​​represent different cache mechanisms.

[0054] Define a dynamic selection algorithm A(S(D,N,F,C)), which determines the optimal strategy based on the current specific value, for example:

[0055]

[0056] Where T is a threshold value, which is set according to the actual application scenario. When the selected synchronization strategy is full synchronization (A(S)=1) and incremental synchronization (A(S)=2), when the system performance is greater than T, full synchronization with simple logic is selected, and the cache of the entire document or task is refreshed; when the system performance is less than or equal to T, incremental synchronization with high efficiency but more complex logic is selected, and only the modified part is sent, so as to ensure the real-time data transmission while saving system resources.

[0057] A method for operating a tag-based task information extraction system is as follows, comprising the following steps:

[0058] S1: Create a document through the document management module, multiple users fill in the document, and the updated content and updated users are displayed in real time on the document;

[0059] S2: Create task labels and generate labels and task attributes based on the updated documents;

[0060] S3: Perform project management based on the generated tags and task attributes.

[0061] Through the above solution, automatic matching of task attributes and document information is achieved, which greatly reduces the risk of errors caused by manual operations and is conducive to improving accuracy; and data synchronization between tasks and documents can be completed quickly, avoiding the problem of synchronization delay in existing methods.

[0062] Project management in S3 can return tasks, and you can view document information by clicking on the task label.

[0063] Project management in S3 includes task management and automated processes.

[0064] A readable storage medium stores a computer program, which, when executed by a processor, implements the method steps for running a tag-based task information extraction system.

[0065] The above solution has the following advantages:

[0066] The association between the document management module and the tag management module is realized through the embedded module of the present invention, which solves the problem of disconnection between tasks and documents, enables tasks to be embedded in documents and kept dynamically updated, thereby realizing the connection between tasks and documents in project management; users can jump to the tag management module to view detailed information with one click through the task tag embedded in the document, and are allowed to directly modify the task attributes through the tag management module, which is conducive to simplifying the operation process between tasks and documents; through the embedded design of task tags and strict identification binding, the uniqueness and relevance of task attributes are ensured, while errors caused by manual input are avoided, and the reliability of collaboration is improved; through the two-way synchronization module, task attributes and document information are synchronized in real time to ensure the consistency and timeliness of information, effectively improve the efficiency of collaboration, and through the dynamic fusion of tasks and documents, the time for task status review, information synchronization and content update is significantly reduced, making project management more intuitive and efficient, meeting the actual needs of complex project collaboration, and helping to simplify operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below according to specific embodiments of the present invention in conjunction with the accompanying drawings, wherein:

[0068] Figure 1 A flowchart of a tag-based task information extraction system, method and medium; DETAILED DESCRIPTION

[0069] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0070] In specific embodiment 1, a tag-based task information extraction system of the present invention includes:

[0071] Tag management module: The tag management module can be used to directly create tags, and edit, assign, update the status and store the created tags. Users can add tags to document information to convert it into a manageable object, and automatically generate relevant fields and task attributes to achieve structured management of tags. Tasks in the tag management module can also be generated by identifying task tags embedded in documents. For tracing the task, the document ID and name can be viewed, and the document page can be jumped to.

[0072] Document management module: The document management module is used to create documents, edit and store the created documents;

[0073] Embedding module: The embedding module is used to connect the document management module with the tag management module, so that the document information and task attributes are dynamically bound. The tags embedded in the document can directly reference the relevant fields and task attributes in the tag management module, where the task attributes include title, tag, and status, to achieve the logical association between the task and the document. Users can jump to the tag management module to view detailed information with one click through the task tags embedded in the document, and can directly modify the task attributes through the tag management module to achieve the mutual jump function. The embedding module supports adding tags and withdrawing tags through the document management module page;

[0074] The algorithm process for implementing the logical association between tasks and documents in the embedded module is:

[0075] The text content in the document and the task-related fields in the tag management module are represented by TF-IDF vectors. After the text content in the document is cleaned, segmented, and standardized, the TF-IDF value of each word is calculated.

[0076] TF represents the frequency of a word t appearing in a document d.

[0077]

[0078] IDF represents the importance of a word t in the entire document set D.

[0079]

[0080] TF-IDF is the product of TF and IDF.

[0081] TF-IDF(t,d,D)=TF(t,d)×IDF(t,D);

[0082] Where t represents a word, d represents a document, and D represents the entire document set;

[0083] Aggregate the TF-IDF values ​​of all words into a vector v,

[0084] v=[TF-IDF(t1,d,D),TF-IDF(t2,d,D)...];

[0085] Calculate the L2 norm of vector v and get the high-dimensional vector Norm(d).

[0086]

[0087] Among them, ∥v∥ 2 is the L2 norm of the vector;

[0088] The text content of the final document is represented as a high-dimensional vector Norm(d);

[0089] Use cosine similarity to measure the similarity between documents and tags, calculate the cosine similarity between documents and all tags, select the most relevant tag as the task tag, and logically associate and bind the document with the tag. The formula is as follows:

[0090]

[0091] Where d is the TF-IDF vector of the document and l is the TF-IDF vector of the tag;

[0092] If no tag meets the association conditions, a new tag is created based on the text content of the document or its summary, and re-associated and bound.

[0093] Two-way synchronization module: The two-way synchronization module uses a unique identifier to synchronize task attributes and document information in real time. When a user updates a task attribute, the operation triggers the real-time synchronization mechanism of the two-way synchronization module. The two-way synchronization module sends the update request to the back-end server through the communication protocol. The back-end server writes the changed data into the cache and pushes it to all participating user sessions through the real-time synchronization mechanism, ensuring that all users can view the changed content in real time and keep the interface consistent. At the same time, when other users update, the changes will be synchronized to the current user interface, realizing two-way synchronization and ensuring that document information and task attributes are updated in real time between multiple users and devices.

[0094] The bidirectional synchronization module implements the synchronization mechanism as follows:

[0095] S(D,N,F,C)=αD+βN+γF+δC+∈;

[0096] Among them, α, β, γ, δ are weight factors, indicating the influence of each parameter on the overall performance, and ∈ represents the contribution of other uncontrollable factors;

[0097] The objective function S(D,N,F,C) represents the overall performance of the system.

[0098] Where D∈[0,1] represents the proportion of data update, N is the number of clients (a positive integer), F∈[0,1] represents the update frequency ratio, and C represents the choice of cache strategy. Different values ​​represent different cache mechanisms.

[0099] Define a dynamic selection algorithm A(S(D,N,F,C)), which determines the optimal strategy based on the current specific value, for example:

[0100]

[0101] Where T is a threshold value, which is set according to the actual application scenario. When the selected synchronization strategy is full synchronization (A(S)=1) and incremental synchronization (A(S)=2), when the system performance is greater than T, full synchronization with simple logic is selected, and the cache of the entire document or task is refreshed; when the system performance is less than or equal to T, incremental synchronization with high efficiency but more complex logic is selected, and only the modified part is sent, so as to ensure the real-time data transmission while saving system resources.

[0102] By embedding the module to associate the document management module with the tag management module, the problem of disconnection between tasks and documents is solved, so that tasks can be embedded in documents and kept dynamically updated, thereby realizing the connection between tasks and documents in project management; users can jump to the tag management module to view detailed information with one click through the task tag embedded in the document, and at the same time allow task properties to be directly modified through the tag management module, which is conducive to simplifying the operation process between tasks and documents; through the embedded design of task tags and strict identification binding, the uniqueness and relevance of task attributes are ensured, while errors caused by manual input are avoided, and the reliability of collaboration is improved; through the two-way synchronization module, task attributes and document information are synchronized in real time to ensure the consistency and timeliness of information, effectively improving collaboration efficiency, and through the dynamic fusion of tasks and documents, the time for task status review, information synchronization and content update is significantly reduced, making project management more intuitive and efficient, meeting the actual needs of complex project collaboration, and helping to simplify operations.

[0103] like Figure 1 As shown, a method for operating a label-based task information extraction system is as follows, comprising the following steps:

[0104] S1: Create a document through the document management module, multiple users fill in the document, and the updated content and updated users are displayed in real time on the document;

[0105] S2: Create task labels and generate tasks based on the updated documents;

[0106] S3: Perform project management based on the generated tasks.

[0107] In the specific embodiment 2, the difference between this embodiment and embodiment 1 is that in this embodiment, the use of a unique identifier to synchronize task attributes and document information in real time can be replaced by a hyperlink-based task tag format, which directly links to the URL address of the corresponding task in the tag management module, eliminating the explicit use of the unique identifier. When the task attribute is updated, the linked API is automatically called to complete the synchronization operation; or the task tag can be replaced by not being displayed in the document, but attached to the paragraph in the form of hidden metadata, such as in HTML. <meta> Tags or embedded properties of Word documents are used to associate and synchronize tasks with documents by reading paragraph metadata.

[0108] In the specific embodiment 3, the difference between this embodiment and embodiments 1 and 2 is that a timed polling module is provided in this embodiment, and the timed polling module is used to periodically send query requests for document information to the tag management module to check whether the task attributes are updated. If there are changes, they are synchronized to the document. The timed polling module is used to replace the two-way synchronization module in embodiment 1; the two-way synchronization module can also be replaced by using the WebSocket protocol to establish a long connection, and the tag management module actively pushes status updates to the document management module, which can achieve more efficient real-time synchronization, but has higher requirements on system resources.

[0109] In the specific embodiment 4, the difference between this embodiment and embodiments 1 to 3 is that in this embodiment, the task tag embedded in the document is replaced by detailed information of the task, such as status, person in charge, etc., which is suspended in the document, so that users can quickly view and edit tasks without jumping; or an embedded task editing panel is implemented in the document, and users can expand the panel by clicking on the task tag and complete the viewing and editing operations of the task directly in the document without leaving the current page.

[0110] In specific embodiment 5, the difference between this embodiment and embodiments 1 to 4 is that in this embodiment, the direct association of tasks with document paragraphs through task tags and unique identifiers can be replaced by the association achieved by mapping document chapter numbers and task numbers. For example, the task number is automatically bound to the chapter number of the document, and the task is matched by parsing the document chapter structure; or it is replaced by automatic matching technology based on semantic analysis, which performs semantic matching between keywords in the task description and the document content, and dynamically establishes the association between tasks and documents without the need to manually insert task tags.

[0111] In the specific embodiment 6, the difference between this embodiment and embodiments 1 to 5 is that in this embodiment, the document management module and the tag management module can be linked to each other through embedded modules, and replaced by a unified document-task collaborative editing platform, in which tasks and documents are treated as the same object, and tasks are directly embedded in the document structure. When the document is saved, the task status is saved synchronously, and all task and document operations are completed in one interface, thereby eliminating the need for synchronization operations between modules.

[0112] When working, first create a document through the document management module, and multiple users fill in the document at the same time. Yjs uses WebSocket technology to achieve real-time data synchronization between multiple clients. The updated content and the name of the user being updated can be displayed in real time on the document. According to the updated document, create a task tag to generate the task. Click the task tag to jump to the document to view the information with one click, which is convenient for viewing or modifying detailed information. You can also modify the task properties directly in the tag management module; finally, carry out project management based on the generated tasks.

[0113] Obviously, the above embodiments are merely examples for clear explanation and are not limitations on the implementation methods. For ordinary technicians in the relevant field, other different forms of changes or modifications can be made on the basis of the above description. It is not necessary and impossible to list all the implementation methods here, and the obvious changes or modifications derived therefrom are still within the protection scope of the invention.

Claims

1. A tag-based task information extraction system, characterized in that: include: The tag management module is used to create tags, and edit, assign, update the status of and store the created tags. Users can add tags to document information to convert it into a manageable object, and automatically generate related fields and task attributes to achieve structured management of tags; A document management module, which is used to create documents, edit and store the created documents; An embedding module is used to connect the document management module with the tag management module, so that the document information and the task attributes are dynamically bound. The tags embedded in the document can directly reference the relevant fields and task attributes in the tag management module to achieve the logical association between the task and the document. The user can jump to the tag management module to view the detailed information with one click through the task tags embedded in the document, thus achieving the mutual jump function. The embedding module supports adding tags and withdrawing tags through the document management module page; Bidirectional synchronization module: When a user updates a task attribute, the operation triggers the real-time synchronization mechanism of the bidirectional synchronization module. The bidirectional synchronization module sends the update request to the back-end server through the communication protocol. The back-end server writes the changed data into the cache and pushes it to all participating user sessions through the real-time synchronization mechanism, ensuring that all users can view the changed content in real time and keep the interface consistent. At the same time, when other users update, the changes will be synchronized to the current user interface, realizing bidirectional synchronization and ensuring that document information and task attributes are updated in real time between multiple users and devices.

2. A tag-based task information extraction system as claimed in claim 1, characterized in that: By identifying the task tags embedded in the document, tasks in the tag management module are generated, and the task is traced back to view the document ID and name, and jump to the document page.

3. A tag-based task information extraction system as claimed in claim 1, characterized in that: The task attributes include title, label, and status.

4. A tag-based task information extraction system as claimed in claim 1, characterized in that: In the mutual jump function, it is allowed to directly modify the task attributes through the tag management module.

5. The tag-based task information extraction system according to claim 1, characterized in that: The bidirectional synchronization module synchronizes task attributes and document information in real time using a unique identifier.

6. A tag-based task information extraction system as claimed in claim 1, characterized in that: The algorithm process for implementing the logical association between tasks and documents in the embedded module is: The text content in the document and the task-related fields in the tag management module are represented by TF-IDF vectors. After the text content in the document is cleaned, segmented, and standardized, the TF-IDF value of each word is calculated. TF represents the frequency of a word t appearing in a document d. IDF represents the importance of a word t in the entire document set D. TF-IDF is the product of TF and IDF. TF-IDF(t,d,D)=TF(t,d)×IDF(t,D); Where t represents a word, d represents a document, and D represents the entire document set; Aggregate the TF-IDF values ​​of all words into a vector v, v=[TF-IDF(t1,d,D),TF-IDF(t2,d,D)...]; Calculate the L2 norm of the vector v to obtain a high-dimensional vector Norm(d), Among them, ∥v∥2 is the L2 norm of the vector; The text content of the final document is represented as a high-dimensional vector Norm(d); Use cosine similarity to measure the similarity between documents and tags, calculate the cosine similarity between documents and all tags, select the most relevant tag as the task tag, and logically associate and bind the document with the tag. The formula is as follows: Where d is the TF-IDF vector of the document and l is the TF-IDF vector of the tag; If no tag meets the association conditions, a new tag is created based on the text content of the document or its summary, and re-associated and bound.

7. The tag-based task information extraction system according to claim 1, characterized in that: The bidirectional synchronization module implements the synchronization mechanism in the following way: S(D,N,F,C)=αD+βN+γF+δC+∈; Among them, α, β, γ, δ are weight factors, indicating the influence of each parameter on the overall performance, and ∈ represents the contribution of other uncontrollable factors; The objective function S(D,N,F,C) represents the overall performance of the system. Where D∈[0,1] represents the proportion of data update, N is the number of clients (a positive integer), F∈[0,1] represents the update frequency ratio, and C represents the choice of cache strategy. Different values ​​represent different cache mechanisms. Define a dynamic selection algorithm A(S(D,N,F,C)), which determines the optimal strategy based on the current specific value, for example: T is a threshold value, which is set according to the actual application scenario. When the selected synchronization strategy is full synchronization A(S)=1 and incremental synchronization A(S)=2, when the system performance is greater than T, full synchronization with simple logic is selected, and the cache of the entire document or task is refreshed. When the system performance is less than or equal to T, incremental synchronization with high efficiency but more complex logic is selected, and only the modified part is sent, so as to ensure the real-time data transmission while saving system resources.

8. The method for operating the tag-based task information extraction system according to claim 1 is as follows, characterized in that: The following steps are involved: S1: Create a document through the document management module, multiple users fill in the document, and the updated content and updated users are displayed in real time on the document; S2: Create task labels and generate labels and task attributes based on the updated documents; S3: Perform project management based on the generated tags and task attributes.

9. A tag-based task information extraction method as claimed in claim 8, characterized in that: Project management in S3 can return tasks, and you can view document information by clicking on the task label; project management in S3 includes task management and automated processes.

10. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method steps of claim 6 are implemented.

Citation Information

Patent Citations

  • Distributed file sharing system and method supporting multiple clients

    CN104301420A

  • Client data updating method and system and medium

    CN113986937A

  • Document labeling method and device, electronic equipment and storage medium

    CN115659969A

  • Data collaborative management method and device supporting multi-task multi-document production

    CN117422054A

  • To-do task processing method and device, terminal and storage medium

    CN118052409A