A file approval method based on NLP
Through NLP-based technology, dynamic classification of approval documents and intelligent allocation of tasks are realized, and the problems of insufficient adaptability of approval processes and difficult to guarantee the fairness of task allocation in the existing technology are solved, thereby improving approval efficiency.
Patent Information
- Application Number
- CN202510317524.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-18
AI Technical Summary
The existing technology lacks dynamic adjustment capabilities in the classification of approval documents and task allocation, resulting in insufficient adaptability of the approval process and difficulty in ensuring the fairness of task allocation. Some approval personnel are in a state of overload or insufficient task for a long time, affecting the overall approval efficiency.
The NLP-based file approval method is adopted, and the file business classification results and semantic feature vectors are generated through word segmentation processing and semantic vector transformation. Combined with the organizational structure and the task situation of the approval personnel, the matching degree indicators and task processing capabilities are calculated, and task allocation is dynamically adjusted.
It realizes the automated classification of the approval process, improves the semantic understanding of the approval documents, ensures the pertinence and fairness of task allocation, reduces approval delays, and optimizes the overall approval efficiency.
Smart Images

Figure CN119850154B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of administrative management, and particularly to a document approval method based on NLP. Background Art
[0002] The prior art relies on fixed rules to classify approval documents, making it difficult to adapt to text changes in different business scenarios. The classification method of approval documents lacks the ability of dynamic adjustment, affecting the adaptability of the approval process. The allocation method of approval tasks mainly relies on manual experience or preset rules, lacking a data-driven dynamic optimization mechanism, making it difficult to ensure the fairness of task allocation. Some approval personnel are in a state of long-term task overload or task shortage, affecting the overall approval efficiency. Therefore, improvements are needed. Summary of the Invention
[0003] The purpose of the present invention is to solve the drawbacks existing in the prior art, and to propose a document approval method based on NLP.
[0004] To achieve the above purpose, the present invention adopts the following technical solutions. A document approval method based on NLP includes the following steps:
[0005] Input an approval document, perform word segmentation on the content of the approval document, calculate the similarity between the word segmentation result and a preset business type dictionary, and generate a document business classification result; based on the document business classification result, perform semantic vector conversion on the keywords in the content of the approval document, calculate the semantic association strength between the keywords, and generate a document semantic feature vector;
[0006] Based on the document semantic feature vector, call the responsibilities of each department and the job responsibilities of personnel in the organizational structure, calculate the matching value between the vector content and the department and personnel responsibilities, and generate a responsibility matching result; based on the responsibility matching result, count the current task quantity and task processing duration of each approval personnel, and calculate and generate an approval personnel matching degree index;
[0007] Based on the approval personnel matching degree index, set a task allocation threshold, screen candidate approval personnel according to the task allocation threshold, and generate a task processing ability index; based on the task processing ability index, combine the task saturation of the candidate approval personnel to generate an approval task allocation sequence;
[0008] Based on the approval task allocation sequence, extract the information of the first approval personnel in the sequence, allocate the approval document to the first approval personnel, record the allocation time and the document identifier, and generate an approval task allocation record.
[0009] Preferably, the step of obtaining the document business classification result is:
[0010] After inputting the approval document, word segmentation is performed through natural language processing, semantic analysis is carried out on the text in the document, the sentences are broken into independent lexical units to form the word segmentation result;
[0011] According to the word segmentation result, calculate the similarity between each word segmentation element and the entries in the preset business type dictionary to obtain the similarity result;
[0012] Based on the similarity result, compare the similarity scores of each business type, and select the one with the highest similarity score as the type identifier of the document to generate the document business classification result.
[0013] Preferably, the steps for obtaining the document semantic feature vector are as follows:
[0014] Based on the document business classification result, analyze the content of the approval document, extract the keywords in the approval document content, screen the keywords, remove redundant, synonymous or interfering items of the keywords, and at the same time perform word form normalization on the keywords to obtain a keyword list;
[0015] According to the keyword list, calculate the semantic association strength between the keywords. The calculation formula is:
[0016] ;
[0017] Among them, represents the semantic association strength between the keywords, represents the keyword appearance frequency in the text, and respectively represent the and the semantic vector values of the keywords, and respectively represent the relative position deviations within the context window of the and the keywords, and represent the semantic dependency distances of the keywords in the text, and represent the dependency weights of the keywords in different syntactic structures, and represent the number of occurrences of the keywords in different text segments;
[0018] Based on the semantic association strength, construct a document semantic feature vector.
[0019] Preferably, the steps for obtaining the responsibility matching result are as follows:
[0020] Based on the semantic feature vector of the document, call the responsibility descriptions of each department in the organizational structure and the job responsibilities of personnel, parse the responsibility description text, extract responsibility keywords, construct a responsibility keyword vector, and obtain a responsibility keyword vector matrix;
[0021] According to the responsibility keyword vector matrix, calculate the matching value between the document semantic feature vector and the responsibility keyword vector. The expression is:
[0022] ;
[0023] where, represents the matching value, is the -dimensional component of the document semantic feature vector, is the -dimensional component of the responsibility keyword vector, and respectively represent the relative distances of the document keyword and the responsibility keyword in the semantic space, is the word frequency value of the th matching word, and represent the position indexes of the document keyword and the responsibility keyword in the syntactic structure, , and represent the matching similarity coefficients between each responsibility description text, represents the dimension number of the document semantic feature vector, represents the number of matching words in the responsibility keyword vector, represents the total number of responsibility description texts;
[0024] Based on the matching value, conduct a matching evaluation on the responsibility descriptions of each department and the job responsibilities of personnel, and select the responsibility item with the highest matching value as the affiliated department and the responsible person to obtain the responsibility matching result.
[0025] Preferably, the steps for obtaining the matching degree index of the approval personnel are as follows:
[0026] Based on the responsibility matching result, count the current task quantity of each approval personnel, calculate the cumulative task processing duration of the tasks already processed by each approval personnel, and at the same time extract the list of tasks to be processed by the approval personnel to obtain the task statistics information of the approval personnel;
[0027] According to the task statistics information of the approval personnel, analyze the task processing mode of each approval personnel, calculate the average number of tasks completed per unit time, and count the current task saturation degree of each approval personnel to obtain the task load index of the approval personnel;
[0028] Based on the task load index of the approvers, calculate the approver matching degree index, and the calculation formula is:
[0029] ;
[0030] Among them, represents the approver matching degree index, is the task processing duration of the tasks completed by the approver, is the current task quantity of the approver, represents the quantity of tasks to be processed by the approver, represents the average completion duration of the historical tasks of the approver, represents the task processing efficiency of the approver in different task types, represents the standard deviation of the task processing time of the approver in different task categories, and represent the task completion stability coefficients of the approver in different time periods, represents the quantity of task records of the approver, represents the quantity of task categories of the approver.
[0031] Preferably, the steps for obtaining the task processing ability index are:
[0032] Based on the approver matching degree index, set a task allocation threshold, screen the matching degree indexes of all approvers, and extract the approvers with matching degree indexes higher than the task allocation threshold to obtain a list of candidate approvers;
[0033] According to the list of candidate approvers, calculate the task completion duration of each candidate approver, and the calculation formula is:
[0034] ;
[0035] Among them, represents the task completion duration of the candidate approver, is the total duration of the tasks completed by the candidate approver in the past, represents the estimated duration of the current task to be processed, and represent the task switching time of the candidate approver, is the standard processing duration of the historical tasks of the candidate approver, and represent the adjustment parameters of the task complexity, 、 and represent the task processing efficiency of the candidate approver in different time periods;
[0036] Based on the task completion duration, evaluate the task processing capabilities of candidate approvers, and combine the task category, processing timeliness, and execution records to obtain task processing capability metrics.
[0037] Preferably, the steps for obtaining the approval task assignment sequence are as follows:
[0038] Based on the task processing capability metrics, extract the task processing capabilities of all candidate approvers, call the task saturation of the candidate approvers, and analyze the current task quantity and task load situation of each candidate approver to obtain candidate approver task status information;
[0039] According to the candidate approver task status information, perform a matching analysis on the task processing capability metrics and task saturation of the candidate approvers, analyze the acceptable task quantity of each candidate approver, and perform a priority ranking on the candidate approvers based on the acceptable task quantity to generate a candidate approver task assignment priority sequence;
[0040] Based on the candidate approver task assignment priority sequence, select candidate approvers in turn according to the execution requirements of the approval tasks, assign the approval tasks in priority order, and record the task assignment order to generate an approval task assignment sequence.
[0041] Preferably, the steps for obtaining the approval task assignment record are as follows:
[0042] Based on the approval task assignment sequence, extract the information of the first approver in the approval task assignment sequence, call the identity information and position information of the approver, and verify the approval task assignment sequence to generate approval task first assignment information;
[0043] According to the approval task first assignment information, match the file identifier of the file to be approved, call the approval file content and assignment time data, and bind the approval file to the first approver, record the assignment situation of the approval file, and generate approval task assignment binding information;
[0044] Based on the approval task assignment binding information, establish an approval task assignment record, store the approval file assignment time, file identifier, and first approver information, and synchronously update the approval task management database.
[0045] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0046] Through the word segmentation processing of the approval document content, the document business classification result is formed, enabling the approval process to be automatically classified based on text features and reducing the dependence on fixed rules. Combining semantic vector conversion and keyword semantic association calculation enhances the semantic understanding ability of the approval content, ensures the accurate matching of the approval document with the organizational structure responsibilities, and avoids approval errors caused by insufficient keyword matching. By calculating the matching value, the approval document is associated with the department responsibilities and job responsibilities, making the assignment of approval tasks more targeted and avoiding difficulties in task execution due to mismatched responsibilities of the approval personnel. Combining the number of tasks and task processing duration of the approval personnel, the approval personnel matching degree index is comprehensively calculated, so that the selection of approval personnel is not only based on responsibility adaptation, but also takes into account the current task load, reducing approval delays caused by workload imbalance. Setting a task assignment threshold, screening candidate approvers, and calculating the task processing ability index make the task assignment more dynamically adjustable. Combining the task saturation of the approval personnel, a sequence of approval task assignments is constructed to ensure the fairness and stability of task assignment, and to avoid long-term task backlogs or insufficient tasks for some personnel. By extracting the information of the first approval personnel in the approval task assignment sequence, the assignment and recording of the approval task are completed, forming a complete approval task tracking system, making the approval document circulation more transparent, and effectively optimizing the scheduling and execution efficiency of the approval task. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a schematic diagram of the steps of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0049] Please refer to Figure 1 , the present invention provides a technical solution, a document approval method based on NLP, including the following steps:
[0050] Input the approval document, perform word segmentation processing on the approval document content, calculate the similarity between the word segmentation result and the preset business type dictionary, and generate a document business classification result; based on the document business classification result, perform semantic vector conversion on the keywords in the approval document content, calculate the semantic association strength between the keywords, and generate a document semantic feature vector;
[0051] Based on the document semantic feature vector, call the descriptions of the responsibilities of each department and the job responsibilities of personnel in the organizational structure, calculate the matching value between the vector content and the department and personnel responsibilities, and generate a responsibility matching result; based on the responsibility matching result, count the current number of tasks and task processing duration of each approval personnel, and calculate and generate an approval personnel matching degree index;
[0052] Based on the reviewer matching degree index, set the task assignment threshold, screen candidates for reviewers according to the task assignment threshold, and generate task processing ability indicators; based on the task processing ability indicators, combine with the task saturation of the candidate reviewers to generate an approval task assignment sequence;
[0053] Based on the approval task assignment sequence, extract the information of the first reviewer in the sequence, assign the approval document to the first reviewer, record the assignment time and file identifier, and generate an approval task assignment record.
[0054] The steps for obtaining the file business classification result are as follows:
[0055] After inputting the approval document, perform word segmentation through natural language processing, conduct semantic analysis on the text in the document, break the sentences into independent lexical units, and form a word segmentation result;
[0056] According to the word segmentation result, calculate the similarity between each word segmentation element and the entries in the preset business type dictionary to obtain a similarity result;
[0057] Based on the similarity result, compare the similarity scores of each business type, and select the one with the highest similarity score as the type identifier of the file to generate the file business classification result.
[0058] Specifically, after obtaining the approval document, first read its text content, retrieve and parse the character sequence line by line for the complete text, perform word segmentation operations on each segment, compare the word forms and word frequencies in the text according to the pre-aggregated industry general word library and common stop word list. Here, the stop word list is screened manually from the accumulated similar approval documents over the years and formed by collecting words that appear frequently but have no actual meaning. Single-occurrence words are also included in the statistics. The minimum occurrence frequency threshold is specified as 1 and inferred based on historical usage and a certain error tolerance. If a large number of invalid word segmentations are found in the subsequent process, this value will be adjusted as appropriate. Then, organize the processed word segmentation result together with information such as character position, part-of-speech tagging, and frequency tagging and separate them with commas. When encountering words with unclear meanings or homographs, determine the more appropriate word segmentation boundary and part of speech based on their relevance in the corpus. Additionally, for common abbreviations such as short names or industry-specific symbols, additional comparison and matching will be carried out. Check and retain the most contextually appropriate form for words that may overlap multiple times in different contexts. If it is necessary to distinguish the term scope of professional terms in a specific scenario, a list containing domain keywords is set in advance. This list is formulated by several professionals in combination with approval specification documents and technical materials and updated periodically. Finally, summarize all the marked words and record them uniformly to form the final word segmentation result.
[0059] After obtaining the word segmentation results, each word segmentation element is first encoded. By converting the word segmentation element into a word vector or using a string similarity calculation method, it is compared with each item in the prepared business type dictionary one by one. The business type dictionary can be uniformly formatted first, with the key entries in the dictionary arranged in alphabetical or pinyin order, and a category label is assigned to each entry. Then, cosine similarity or character edit distance is selected as the main comparison index. At this time, a similarity threshold can be set to distinguish results with low matching degrees. The value of this threshold can be estimated by the median of the matching accuracies of previously collected industry texts. For example, in recent sample statistics, it is found that when the similarity is lower than 0.6, a high error rate will be brought. Therefore, 0.6 can be used as the initial setting. If it is found that the error is still high in subsequent verification, the threshold can be adjusted again. When comparing each word segmentation element with the dictionary entries, additional weight corrections can also be made according to the word length and the difference between simplified and traditional Chinese characters of Chinese characters. For example, professional terms with more characters are regarded as having a higher priority. Then, combined with the stop word list generated in the early stage, word segmentations that appear frequently but do not correspond to specific business meanings are removed. After calculating the similarity, the matching degree of each word segmentation element corresponding to the dictionary entry is recorded, and a set of corresponding similarity scores is generated for all word segmentation elements. Finally, these score sets are integrated into a similarity result.
[0060] After obtaining the similarity result, the similarity scores of each business type included in it are summarized and compared. First, the similarity values corresponding to all word segmentation elements are weighted and combined to calculate the cumulative similarity of each business type in the overall text. A difference threshold can be added to judge the significant difference degree between scores. This difference threshold can be determined by statistically analyzing the distribution of actual judgment results in historical approval documents. For example, in past data, if the difference between the highest score and the second highest score is less than 0.05, there is often judgment confusion. Then, 0.05 can be selected as the initial reference value. If it is found that the tie situation is still frequent in subsequent tests, the threshold can be continuously adjusted. After confirming the sorting of the similarity scores, the business type ranked first is taken as the type identifier of the document. If the scores of multiple business types are simultaneously higher than a certain set value and meet the limit of the minimum difference range, multi-label marking can also be made in the record. Finally, the obtained document type information is sorted into the document business classification result.
[0061] The steps to obtain the document semantic feature vector are as follows:
[0062] Based on the document business classification result, parse the content of the approval document, extract the keywords in the content of the approval document, screen the keywords, remove redundant, synonymous or interfering items of the keywords, and at the same time normalize the word forms of the keywords to obtain a keyword list;
[0063] According to the keyword list, calculate the semantic association strength between the keywords. The calculation formula is:
[0064] ;
[0065] in, Represents the strength of semantic association between keywords. Representative keywords The frequency of occurrence in the text, and Respectively represent and The semantic vector value of keywords, and Respectively represent and The relative position deviation of keywords in the context window, and Represents the semantic dependency distance of keywords in the text, and Represents the dependency weight of keywords in different syntactic structures, and Represents the number of times a keyword appears in different text fragments;
[0066] Based on the semantic association strength, a file semantic feature vector is constructed.
[0067] Specifically, based on the document business classification results obtained previously, the text of the current approval document content is read line by line and scanned in character sequence. Then, all the words that appear are retrieved and compared with the existing synonym comparison table and basic terminology dictionary. The words that are marked as synonymous or only appear in specific fields are merged. If the cumulative number of appearances of some short words in the entire text is less than 3 and their meanings do not have conceptual reference, they can be identified as interference items and removed to reduce the impact of irrelevant content. During this period, the word form changes containing common roots will be checked and unified. For example, plural forms or mixed case English words will be processed, and redundancy will be controlled through character segmentation and word form restoration. In order to accurately distinguish For words with multiple meanings, the judgment will be made with reference to the position range of the word in the context and the high-frequency collocation information that has been collected. If the word is frequently associated with specific field terms or numerical indicators in the adjacent text range, it will be listed as a reserved category and recorded with its core meaning label. In order to define "low-frequency words" or "rare words", a corresponding frequency threshold will be set. The initial value of the threshold can be 2 and fine-tuned in combination with the text length distribution of historical approval documents. If it is found that there are still many residual low-frequency interference items in subsequent applications, it may be necessary to further lower or increase the threshold. Finally, all processed valid words are normalized to maintain their text consistency in different scenarios to obtain a keyword list.
[0068] The advantage of the formula is that it incorporates multiple-dimensional elements such as the keyword occurrence frequency, the deviation of the context position, the semantic dependency distance, and the dependency weight in the syntactic structure into the operation, so that the association strength not only depends on the number of occurrences of the keyword itself, but is also closely related to its relative distribution and structural relevance in the text.
[0069] The steps to obtain the parameters are as follows:
[0070] First, count the number of occurrences of each keyword in the approval text, compare this number of occurrences with the total number of words in the current text, and then calculate the keyword proportion to form a frequency value. At this time, the frequency value can be written as , where QRR is the number of occurrences of the keyword, and AL is the total number of words in the text. If an example with a total number of words of about 2000 is extracted from the registered approval text and the number of occurrences of this keyword is 50, then the corresponding is . In large-scale texts, the statistical results of multiple segments can also be integrated to form a weighted frequency and then normalized.
[0071] and The steps to obtain the parameters are as follows:
[0072] It is necessary to vectorize each keyword, and the corresponding vector value can be obtained through the previously constructed word vector training. The specific method is to collect several words with known semantic labels from the approval texts of the same type over the years and perform vector training. The training model can adopt the method of fixed-window scanning. For each word, several words within its context range are evenly extracted to construct training samples, and then the maximum likelihood estimation is used to continuously update the word vector parameters. Suppose that when the model finally converges, the vector value corresponding to the keyword "supply chain" is (0.12, 0.35, -0.09), and the vector value corresponding to the keyword "logistics" is (0.10, 0.30, -0.05). At this time can take (0.12, 0.35, -0.09), can take (0.10, 0.30, -0.05), and the same method is used to obtain other keywords.
[0073] and The steps to obtain the parameters are as follows:
[0074] Record the context positions where each keyword appears in the text, and set a context window. For example, define the window span as 5 words, and count the average position coordinates of the keyword in each occurrence within the range of 5 words on each side. Then calculate the difference between this average position coordinate and the starting position of the sentence where the keyword is located, and map the result to the range from -1 to 1. Suppose the average coordinate of the context window where the keyword "project" appears is 12, with the start of the sentence as 0, then we can get , where ZC is the total number of words in the sentence. If the total number of words in the sentence is 30, then .
[0075] and The steps for obtaining the parameters are as follows:
[0076] First, record the distance difference between the occurrences of each keyword between sentences or paragraphs. The distance can be defined as the word count difference or symbol count difference between keywords, or it can be obtained by calculating the number of jumps from node m to node n in the dependency tree. If it is observed that the keywords "purchase" and "production" are separated by 3 edges in the dependency tree, then we can or be assigned a value of 3. If this distance becomes 2 or 4 in other text paragraphs, corresponding records are made for it.
[0077] and The steps for obtaining the parameters are as follows:
[0078] Determine the position properties of the keywords in the grammatical structure respectively, such as subject, predicate, object, etc. For keywords in the subject position, a higher dependency weight can be assigned, while those in the modifying position in a clause can be regarded as having a lower weight. When quantifying these weight values, they can be selected within the range of 0.5 to 2.0 and calibrated in combination with statistical experience. If it is observed in a large number of approval texts that keywords in the subject position are more inclined to be related to the global semantics, then the weight of this subject position can be set to 1.8, and the weight of the modifier position can be set to 0.9. Finally, when calculating the ratio of the dependency weights of two keywords , these actually measured values can be substituted to complete the comparison.
[0079] and The steps for obtaining the parameters are as follows:
[0080] It is necessary to identify the number of occurrences of keywords in different text segments. The approval documents are divided into several segments through previous paragraph segmentation, for example, every 500 words or processed separately for each natural paragraph, and then the frequencies of each keyword in different segments are summarized. If a keyword appears in 10 segments, it is recorded as 10 occurrences. If the number of occurrences in 3 of these segments is significantly higher, a differential counting logic can be additionally used to further label the density of these keywords in local paragraphs. If in a specific example, the keyword "budget" appears in 8 paragraphs and the keyword "contract" appears in 12 paragraphs, then correspondingly = 8、 = 12。
[0081] Calculation process:
[0082] In an approval text scenario, a pair of keywords A and keyword B have the following parameters:
[0083] , , , , ; denominator ;
[0084] Substitute these values into the numerator operation:
[0085] Numerator = ;
[0086] First do multiplication: ;
[0087] Then accumulate: ;
[0088] Denominator = 20;
[0089] Therefore: ;
[0090] This result indicates that when the of this pair of keywords is not very large, and after comprehensively considering factors such as the semantic vector difference, context position deviation, and dependency weight ratio, the final association strength is at a level of approximately 0.1103. If there are other keyword combinations in the same text whose value exceeds 0.60, it indicates that their correlation is more obvious, while the currently calculated value of 0.1103 indicates a weak association between keyword A and keyword B.
[0091] Previously, the semantic association strength between keywords was calculated and After the numerical sequence, collect the correlation strength data corresponding to all keyword combinations, and conduct a distribution statistics on the correlation strength values within the range of 0 to 1. If the correlation strength of some keyword pairs exceeds 0.70, they can be classified into the high-correlation group. If it is between 0.30 and 0.70, it is marked as the medium-correlation range. If it is below 0.30, it is considered weakly correlated. To refine its applicability in subsequent text descriptions, the keywords belonging to the high-correlation group need to be aggregated and recorded under the same keyword cluster. Here, the correlation values can be compared again within the high-correlation group by pairwise comparison. If both are greater than 0.60, they are classified into the same cluster. Otherwise, when it is greater than 0.30 and less than 0.60, it is tentatively designated as a neighboring cluster. Using such a hierarchical grouping method, the theme of the large-scale approval text can be mapped at the keyword level. Then, compare one by one with the main type labels in the business classification results to see if there is a keyword cluster that corresponds to a specific business category. If it is determined that there is a high overlap, the keyword cluster can be bound to the category for subsequent traceability and retrieval. Finally, summarize all keywords and their cluster affiliations, and construct a file semantic feature vector with the support of the correlation strength results.
[0092] The steps to obtain the responsibility matching result are as follows:
[0093] Based on the file semantic feature vector, call the responsibility descriptions of each department in the organizational structure and the job responsibility descriptions of personnel, parse the responsibility description text, extract responsibility keywords, construct a responsibility keyword vector, and obtain a responsibility keyword vector matrix;
[0094] According to the responsibility keyword vector matrix, calculate the matching value between the file semantic feature vector and the responsibility keyword vector. The expression is:
[0095] ;
[0096] Among them, represents the matching value, is the dimensional component of the file semantic feature vector, is the dimensional component of the responsibility keyword vector, and respectively represent the relative distances of the file keyword and the responsibility keyword in the semantic space, is the word frequency value of the and represent the position indexes of the file keyword and the responsibility keyword in the syntactic structure, , and represent the matching similarity coefficients between each responsibility description text, represents the dimensionality of the file semantic feature vector, represents the number of matching words in the duty keyword vector, represents the total number of duty description texts;
[0097] Based on the matching value, the duty descriptions of each department and the job responsibilities of personnel are matched and evaluated, and the duty item with the highest matching value is selected as the affiliated department and the responsible person to obtain the duty matching result.
[0098] Specifically, based on the file semantic feature vector, first read the text content of the duty descriptions of each department and the job responsibilities of personnel in the organizational structure, and determine the proper nouns, job titles, and function description items in the text through the method of continuous sentence-by-sentence retrieval. Then, extract the representative words from these contents one by one and merge them. The duty expressions with synonymous or highly overlapping expressions will be merged. Here, the duty content text will be divided into several small paragraphs, each paragraph covering a range of about 200 to 300 words, and check whether the key verbs and nouns appearing in it are included in the job responsibility word list constructed before. When it is found that some words appear more than 3 times with the common duty or authority description words in the job responsibility word list, they will be marked as possible duty keywords, and the paragraph number where the corresponding words are located and the text distance from other possible duty keywords will be recorded. The text distance can be statistically accumulated by the number of characters or words. Each time a verb with a job-directing meaning such as "review" or "approve" is encountered, the distance from the current keyword will be checked. When the distance reaches the range of 50 to 80 words, a far distribution will be recorded, and when it is within the range of 50 words, a near distribution will be recorded. By comparing the multiple distribution results, the tightness between the keyword and the job description is judged. When the text distance value distribution tends to be concentrated in a small range, it indicates a high degree of correlation. At the same time, semantic recognition is performed on the keywords with a frequency of more than 2 times. When referring to the organizational structure information, it can be found whether the keyword corresponds to the department duty field. If some keywords have a high degree of correlation with multiple departments, they will be further subdivided and compared with other description details in the job responsibility description text, such as whether they contain specific authority scopes or specific work scenarios. Finally, a set of duty keywords that are redundant-free and dominant in the job description text is obtained, and these duty keywords are numbered in the order of text records. When converting them into vector form one by one based on the numbering order, the previously completed vector training method will be referred to, and the corresponding job components or authority components in the department information will be superimposed to form multiple dimensions of the duty keyword vector. The vectors corresponding to all keywords are grouped according to the job or department category and stored in a unified data structure according to the collected text distribution situation. After merging and sorting, the duty keyword vector matrix is obtained.
[0099] The advantage of the formula is that it incorporates into the comprehensive operation process the differences in multiple dimensions, word frequencies, position indexes, and similarity coefficients of each responsibility description text between the file semantic feature vector and the responsibility keyword vector, combines information such as vector dimension differences, word frequency ratios, syntactic position differences, and cross-text similarities, and then obtains a more targeted matching value.
[0100] The steps for obtaining the parameter are as follows:
[0101] This parameter is the component of the file semantic feature vector in the dimension, which is specifically obtained through the text semantic vector training process mentioned above. When training the vector, high-frequency keywords and phrases are extracted from a large number of approval documents, their context dependencies and context distributions are learned, and each keyword is given a vector representation in several dimensions. Subsequently, the keyword vectors contained in the entire approval document are aggregated through weighting or other forms to generate the final file semantic feature vector. is the specific value in the th dimension. This value is determined based on the density, occurrence range, and relevance to the context of the keyword in the text. For example, 500 approved documents are collected, a semantic space with a word vector dimension of 100 is constructed, the keyword set obtained previously is selected, and aggregation is performed one by one. During the aggregation process, keywords with an occurrence frequency greater than 3 and a relevance to the text theme reaching more than 60% are given higher weights, and finally, the file semantic feature vector is synthesized. Assume that the value of the 10th dimension is approximately 0.27 after statistics and normalization.
[0102] The steps for obtaining the parameter are as follows:
[0103] This parameter represents the value of the responsibility keyword vector in the dimension, which is obtained based on the same vector training principle. However, the selected corpus all comes from the texts of department responsibility descriptions and job responsibility descriptions. For the core keywords in each responsibility text, their occurrence frequencies are first recorded, and then a dedicated responsibility keyword vector is generated using the same dimension aggregation idea. is the specific value of this vector in the dimension. To make the file semantic feature vector and the responsibility keyword vector comparable, their training schemes need to be consistent, so that corresponding semantic components can be obtained in the same dimension. For example, when analyzing the department responsibility description of the human resources department, statistics are made on keywords such as "recruitment", "attendance", and "performance", and they are matched with the previously set training model to generate a vector, which may obtain a 100-dimensional vector representation. Among them, the 10th dimension is 0.20. At this time = 0.20.
[0104] With The steps for obtaining the parameters are as follows:
[0105] These two parameters represent the relative distances of file keywords and responsibility keywords in the semantic space. It is necessary to extract the vector coordinates corresponding to the keywords and calculate the Euclidean distance or cosine angle between them in the vector space, and then perform appropriate normalization processing to map the results to a range from 0 to 1. The larger the value, the farther the distance between the two in the semantic space; the smaller the value, the closer they are. When collecting a batch of common department responsibility keywords, it may be found that the distance between "audit" and "approval" in the vector space is relatively small. Therefore it may be close to 0.10, while the distance of keywords related to "market" may be 0.35, etc. After measuring the distance for each keyword pair, the corresponding and are used for the difference operation in the formula. For example, select keyword pairs from 200 responsibility description texts and corresponding approval documents, and calculate the distances one by one. It is found that a certain distance is 0.25. After normalization, the final value can be taken as 0.25 as , and for another distance of 0.30, it can be taken as .
[0106] The steps for obtaining the parameters are as follows:
[0107] This parameter represents the word frequency value of the th matching word in the responsibility text, which is statistically calculated based on the total number of words in the department responsibility description and the job responsibility description and the number of times the matching word appears. For example, if the key responsibility word "finance" appears 150 times in a 30,000-word responsibility text, its word frequency value can be initially recorded as , and then further segmented in combination with the corresponding job level or sentence weight coefficient. In order to obtain a more practically meaningful value, this result can be compared with other key responsibility words and normalized. If it is finally calculated that the word frequency proportion of "finance" in the whole text ranks in the top 10%, can be given a weight value of 0.90. If it ranks in the bottom 10%, a weight value of 0.30 can be considered. Finally, these word frequency values are brought into the formula for product operation. The example of obtaining here is: when the number of times the matching word "compliance" appears is 120 and the length of the text where it is located is about 24,000 words, its preliminary word frequency is 0.005. If subsequent comparison shows that its appearance frequency ranks in the 30% position among similar keywords, it will finally be recorded as = 0.70.
[0108] and The steps for obtaining the parameters are as follows:
[0109] These two parameters are used to represent the position indexes of the document keyword and the responsibility keyword in the syntactic structure. First, the corresponding sentence needs to be split according to the word order, and a fixed sequential position is set for each word in it. Then, the index numbers where the document keyword and the responsibility keyword appear are recorded respectively. When the syntactic structure is relatively complex, a dependency tree or a sentence analysis tree can be built first, and the nodes are traversed in turn to obtain the keyword positions, and then the position values are mapped to a reasonable range. For example, in a sentence with 40 words, if the document keyword appears in the 5th word, then = 5, and if the responsibility keyword appears in the 10th word, then = 10. Then, a normalization in the range of 0 to 1 is performed. For example, map 5 to , and map 10 to etc. Finally, it is stored for subsequent calculations. If in a longer description, the index numbers can be proportionally allocated in a larger range. Taking an actual example: a certain sentence has 50 words, the document keyword appears in the 8th word, and the responsibility keyword appears in the 20th word. After normalization , .
[0110] 、 、 The steps for obtaining the parameters are as follows:
[0111] These three parameters respectively represent the matching similarity coefficients between each responsibility description text and the key content of the document, which are used to measure the similarity degree extracted from multiple departments or multiple job descriptions. It is necessary to compare the key paragraphs of different responsibility texts with the keywords in the document. The word vector model constructed in the previous corpus training can be selected to measure the cosine similarity between the sentence or paragraph vectors in each responsibility text and the document vector. Then, the results are statistically analyzed by segments and finally mapped to the range of 0 to 1. For example, when analyzing the responsibility text of Department A, if the average similarity obtained is 0.73 in total, then can be assigned 0.73. If the average similarity obtained for the responsibility text of Department B is 0.45, then is taken as 0.45. If the similarity obtained for the text of Department C is 0.67, then record = 0.67. These values form a list in the multi-department and multi-text scenario for the cumulative operation of the denominator term of the formula to distinguish the overall matching degree of different responsibility texts with the current document. For example, when a certain approval document is compared with the responsibility texts of 5 different departments in turn, and the similarity values obtained are 0.70, 0.63, 0.54, 0.77, 0.58 respectively, then = 0.70, = 0.63, = 0.54… and so on are recorded and substituted into the calculation.
[0112] The steps for obtaining the parameter are as follows:
[0113] This parameter represents the dimensionality of the document semantic feature vector, which is usually fixed during the corpus training phase. For example, if 100 dimensions or 150 dimensions are selected when constructing word vectors, then it corresponds to this determined value. When using 100-dimensional word vectors, the entire document semantic feature vector is a projection in this 100-dimensional space. Therefore, = 100. To obtain a specific and feasible dimension, multiple experiments need to be conducted on the previously collected industry texts. Comparisons can be made between dimensions ranging from 50 to 200 to observe the matching accuracy and computational consumption at each dimension. Then, a dimension number is finally selected on the premise of balancing accuracy and efficiency. For example, after monitoring 120 simulation tests, if the 100-dimensional word vectors show better effects, then = 100 is substituted into the relevant calculations.
[0114] The steps for obtaining the parameter are as follows:
[0115] This parameter is the number of matching words in the responsibility keyword vector. A series of keywords that play a major role in the responsibility text need to be retrieved first, and they are counted when constructing the responsibility keyword vector matrix. For example, 20 high-frequency or high-weight keywords are retrieved from the responsibility description of a certain department. After intersecting and comparing with the document keyword list, only 10 of them may actually be matched. Thus, = 10. When the retrieval scope is expanded or the responsibilities of multiple positions are combined, it will also change accordingly. For example, when combining 50 keywords from the texts of three positions and screening out 25 keywords that match the document, then = 25. Finally, these values are substituted into the corresponding formula for subsequent multiplication operations.
[0116] The steps for obtaining the parameter are as follows:
[0117] This parameter represents the total number of responsibility description texts. When analyzing an organizational structure, the responsibility descriptions of each department need to be considered comprehensively, and then each text is counted separately. In some scenarios, there are 10 departments, and each department such as human resources, finance, logistics, and sales has one or more responsibility documents. Then, after all are combined, there may be a total of 15 to 20 texts. At this time, is this total number. For example, when it is determined through organizational structure inspection that there are 18 responsibility description documents for positions, then = 18, during the statistical stage, these texts will be successively subjected to keyword parsing and vector construction, and then participate in the combined calculation of the denominator part of this formula.
[0118] Calculation process:
[0119] When , , , and after obtaining the relevant parameters in the front, substitute them into the formula, then:
[0120] ;
[0121] ;
[0122] ;
[0123] Then the numerator part = ;
[0124] The denominator part = ;
[0125] So:
[0126] ;
[0127] Since the normalization interval mentioned above is from 0 to 1, when the calculated value exceeds 1.0, it can be marked as 1.0 in the matching evaluation to reflect a higher duty matching degree.
[0128] The result shows that the matching value obtained in this example is greater than 1.0, indicating that there are relatively close characteristics in terms of dimensional component difference, keyword position index, cross - text similarity, etc. between the file and the selected duty description text, meaning there is a high overlap in the duty content. If the actual normalization results in = 1.0, it means the highest matching degree, and this duty can be regarded as the priority attribution item for matching evaluation in the subsequent steps.
[0129] Based on the matching values, first retrieve the corresponding department responsibility descriptions according to the semantic feature vector information of the acquired documents. Divide the matching values into different intervals and record them item by item. When the matching value is greater than 0.75, it is regarded as a high matching degree, and it can be marked and summarized into a candidate set. When the matching value is between 0.40 and 0.75, it is regarded as a medium matching degree, and it is also marked and stored uniformly. If the matching value is lower than 0.40, it is recorded as a relatively weak match and directly appended to the corresponding department or position name in numerical form. Subsequently, the items with relatively high matching values are centrally screened again. Here, the distribution of the matching values can be compared with the focus of the department responsibility descriptions. For example, when it is found that most of the keywords in the same position responsibility text have a high degree of coincidence with the key themes of the document, this text will be regarded as a more suitable option. Then, compare several position descriptions with a matching degree of more than 0.80 one by one to see if they involve specific management categories or professional categories in the document theme. If the number of completely docked categories reaches 3 or more, this position responsibility will be put into the final list to be selected. During this period, the hierarchical relationship in the organizational structure will also be compared layer by layer, sort all the position texts and mark their matching value scores, list them from the highest to the lowest in order, and then extract the position with the highest matching value first from the list and associate it with its affiliated department to determine the affiliated department and the responsible person. Finally, merge the results into a centralized record to obtain the responsibility matching result.
[0130] The steps to obtain the matching degree index of the approvers are as follows:
[0131] Based on the responsibility matching result, count the current task quantity of each approver, calculate the cumulative task processing duration of the tasks already processed by each approver, and at the same time extract the list of tasks to be processed by the approvers to obtain the task statistical information of the approvers;
[0132] According to the task statistical information of the approvers, analyze the task processing mode of each approver, calculate the average number of tasks completed per unit time, and count the current task saturation degree of each approver to obtain the task load index of the approvers;
[0133] Based on the task load index of the approvers, calculate the matching degree index of the approvers. The calculation formula is:
[0134] ;
[0135] Among them, represents the matching degree index of the approvers, is the task processing duration of the tasks already completed by the approvers, is the current task quantity of the approvers, represents the quantity of tasks to be processed by the approvers, represents the average completion duration of the historical tasks of the approvers, Represent the task processing efficiency of approval personnel on different task types, Represent the standard deviation of task processing time of approval personnel on different task categories, and Represent the task completion stability coefficient of approval personnel in different time periods, Represent the number of task records of approval personnel, Represent the number of task categories of approval personnel.
[0136] Specifically, based on the responsibility matching results, first retrieve the completed task records from the work information database of each approval personnel. By comparing the identification numbers and completion times of tasks one by one, count the current task quantity of each approval personnel. Then, accumulate and summarize according to the actual execution duration of each task. Here, the duration of each completed task will be extracted, and its start time and end time will be recorded. The processing duration of a single task is obtained by subtraction. Then, add up the durations of all completed tasks of the same approval personnel to get the cumulative value. To ensure the accuracy of the duration data, it is necessary to compare these records with the clock-in or operation traces in the log system to confirm whether the corresponding task is associated with the approval personnel. If duplicate registrations or task numbers transferred across departments are found during the retrieval, it is necessary to judge the start and end points of this part of the records in different cases, calculate the duration in a unified format and then accumulate it. During this period, the list of unfinished tasks will also be included in the statistics. For the tasks to be processed by each approval personnel, status such as "unfinished" or "processing" will be marked in the system, and at the same time, the estimated processing time limit required for the task will be recorded. If the processing time limit exceeds the benchmark value set in advance by the department, for example, there is still no result after more than 48 hours, a reminder mechanism may be triggered. The reason for choosing 48 hours as the benchmark value is that in nearly 100 approval cases, the average required duration of most approval processes is about 36 hours, and an additional 12 hours of margin is given to prevent delays caused by multi-department collaboration factors. When recording these tasks to be processed, it is also necessary to check their task categories and priority levels. If the priority level is marked high, it may be listed and sorted separately later. Finally, merge the cumulative duration of all completed tasks, the current list of tasks to be processed that are not completed, and the task quantity to form centralized data to obtain the task statistics information of approval personnel.
[0137] According to the task statistics information of the approvers, it is necessary to analyze the past processing cycle and task structure of each approver. By calculating the number of tasks completed by each approver within a fixed period, the average number of tasks completed per unit time can be obtained. Then, comparing this average value with the current task list in progress, observing the corresponding relationship between the total number of actual tasks in progress and the average processing rate per person. Here, a standard working week can be defined as 40 hours, and the overtime frequency is determined according to the department's habits or enterprise regulations. For example, some departments do not handle tasks on weekends, while others have a 24-hour duty system. Therefore, when calculating the processing efficiency per unit time, the available time periods should be distinguished according to the department's work and rest conditions. Subsequently, comparing the processing efficiency with the current task quantity. If an approver has 8 current tasks, and considering that the average processing rate of this person in the past two weeks is 0.2 tasks per hour, that is, about 8 tasks can be completed in a week, it can be determined that the task saturation degree of this approver tends to be critical. At this time, if new approval tasks are assigned, it may cause queuing or delays. If someone only has 2 tasks backlogged or has a higher processing rate per unit time under the same conditions, it means that their current work saturation is relatively lower. In addition to the speed index, the differences in task types should also be checked. If an approver currently has all high-priority tasks involving cross-departmental signatures, an additional identifier can be given during the statistics to avoid simply equating the number of such tasks with those of low-priority tasks with loose time limits. If it is found during the retrieval that someone's work saturation degree exceeds the threshold set by the department, for example, more than 80% of the working hours are occupied by high-priority tasks, it is recorded as a high-saturation state. The specific value of this threshold can be obtained from the load statistics of nearly a hundred approval processes. By observing the actual operation efficiency of the approval process in segments and continuously testing between 70% and 90%, the most appropriate 80% interval is obtained. Finally, the saturation calculation results of all approvers are integrated into the approver task load index.
[0138] The benefit of the formula is that when evaluating the approver matching degree, it not only integrates the processing duration of the completed tasks and the current task quantity, but also introduces multiple factors such as the quantity of tasks to be processed, the historical average duration, the efficiency differences of different task types, and the task completion stability. Through the quantitative processing of multiple parameters and the coordinated ratio of the numerator and denominator, it provides more refined data support for the subsequent automatic assignment of approval tasks.
[0139] The steps to obtain the parameters are as follows:
[0140] This parameter represents the task processing duration of tasks completed by the approver. When counting each completed task, the start time and end time need to be read from the system log, and the processing duration of the task is accurately obtained by the difference between the two. Then, the durations of all completed tasks of the same approver are recorded in sequence according to the task number. These records form a duration array, and each element corresponds to a task duration. Subsequently, when calculating a direct addition operation will be performed on the elements in this array. To make the values comparable, invalid or interrupted tasks need to be excluded after the preliminary statistics. For example, in a few large-scale cross-month collaborative tasks, the actual time invested by the approver may be segmented, and the durations of multiple segments need to be added together and then combined into the final duration of the task. Finally, a set of accurate completed task duration sets is obtained. For example, by monitoring 5 closed tasks within a month and recording that they consumed 2.5 hours, 3.0 hours, 1.8 hours, 2.0 hours, and 4.5 hours respectively, then the corresponding hours. The example of obtaining here is as described above. Anyone can obtain the specific values of each item by checking the historical task list of this approver and accumulating the start-end time differences.
[0141] The steps to obtain the parameter are as follows:
[0142] This parameter is the details of the current number of tasks of the approver. First, check all the unfinished task lists of this person in the current state, and then record the numbers of tasks of different types or priorities separately. For example, there are 7 general approval tasks, 2 urgent approval tasks, and 1 cross-departmental collaborative task. Then, these 10 tasks can be counted as a unified current task number list, or the category numbers of the tasks can be marked internally, such as 0 for general type, 1 for urgent type, 2 for cross-departmental type, etc. When performing a multiplication operation, such as corresponding to three categories from 1 to 3 respectively, and the specific values are read from this list. For example, when an approver finds 7 general type tasks, 2 urgent type tasks, and 1 cross-departmental collaborative task during statistics, then 、 、 can be obtained. If quantification or power delegation is carried out for different categories, it can also be balanced and expressed through other parameters in the subsequent formula. Example of obtaining: Retrieve the person's daily task list from the background database and compare the status tags one by one to confirm that there are 10 tasks in total of 3 categories in the tasks in progress. Then, allocate them item by item to and other positions.
[0143] The steps to obtain the parameter are as follows:
[0144] This parameter represents the number of tasks awaiting approval by the approver. It needs to be distinguished from the difference in the previous step of "current tasks in progress". "Tasks awaiting approval" focuses more on those tasks that have not yet officially entered the processing stage but have been assigned to the approver by the system. If there are still several unopened or unconfirmed tasks in the to-do list of a certain person in the system, the quantities of these tasks should be accumulated to form , and it is checked against the assignment information in the database. If it is found that the task has been shared and assigned to multiple people during multi-department collaboration, it can only be included in the personal tasks awaiting approval of this approver after the approver confirms receiving the order, so as to avoid double counting. In actual statistics, it can be seen that the to-do lists of some approvers may contain 2 to 5 tasks awaiting approval. If it exceeds 5, there is often short-term congestion. For example: When querying the to-do list of a person in the system, it is found that 4 records among them have the status of "unconfirmed click", and another 2 are in the status of "requiring joint signature" and have not officially started. This person has been assigned by the system, so the total is .
[0145] The steps to obtain the parameter are as follows:
[0146] This parameter represents the average completion duration of the approver's historical tasks. It is necessary to extract the processing times of several completed tasks from their historical data and then perform an average operation. To make the result representative, historical samples of the past three months or six months can be selected, and extremely short or extremely long tasks that are significantly abnormal are excluded. After that, the processing durations of all tasks included in the statistics are added up and divided by the total number of tasks to obtain the average completion duration. If 20 completed tasks were collected in the past three months and their total processing duration is 60 hours, then hours. For example: Randomly select the task records of multiple approvers within a department and perform this mean statistics for each person. If a person finally obtains an average completion duration of 3.2 hours, then 3.2 is used as to be written into the subsequent formula.
[0147] The steps to obtain the parameter are as follows:
[0148] This parameter is used to reflect the task processing efficiency of the approver in different task types. The completion times of different types such as ordinary tasks, urgent tasks, and complex collaboration tasks can be statistically analyzed separately, and then combined with the outputs of these types to compare with the average level of the same type of tasks to form a rate or efficiency score. For example, on average, 0.4 pieces are completed per hour for ordinary tasks and 0.2 pieces are completed per hour for urgent tasks. These sub-items can be combined, and the final value can be obtained by using the method of weighted or comparing with the industry benchmark of the same position. To make To be directly used by the formula, it needs to be mapped to the range of 0 to 1. For example, if it is found through comprehensive evaluation that someone's efficiency is in the top 20% of all personnel, a value of 0.80 can be assigned; if it is in the bottom 20%, a value of 0.20 can be assigned. Example: Statistically analyze the completion rate of various types of tasks of a certain approval personnel in the past month. After comparing it with the department average, it is determined that their efficiency coefficient is 0.75, that is 。
[0149] The steps to obtain the parameter are as follows:
[0150] This parameter represents the standard deviation of the task processing time of the approval personnel for different task categories. It is necessary to list one by one the completion durations of each type of task, then calculate the standard deviation corresponding to each type of task, and finally merge them into a unified value. Here, the root mean square or weighted mean method can be used during the merging to synthesize the standard deviations of different categories. If the standard deviation of the general approval records is 1.2 hours and that of the urgent approval records is 1.8 hours, then it may be between 1.3 and 1.5 hours. Through normalization, it can also be mapped to the range of 0 to 1. For example, 1.3 hours corresponds to 0.65, 1.5 hours corresponds to 0.75, etc. Example: The processing durations of 10 general tasks of a person in the past fluctuated between 3 and 5 hours, and the calculated standard deviation was 0.7 hours. Another 5 urgent tasks had a greater fluctuation, with a standard deviation of 1.5 hours. After summarizing the two according to the task volume weight, it may be 1.1 hours, and then mapped to 0.55, that is 。
[0151] and The steps to obtain the parameter are as follows:
[0152] These two parameters characterize the task completion stability coefficient of the approval personnel at different time periods. It is necessary to quantify the task completion performance of this person during different time periods such as weekdays and non - weekdays, or during the day and at night. It may be through statistics of the total number of tasks, completion efficiency, delay situations, etc. within the time period, and then convert the results into a stability index between 0 and 1. If a stable processing rate can be maintained for most of the working days, then it will be relatively high, for example, between 0.80 and 0.90. If the processing efficiency fluctuates significantly at night, then it will be relatively low, for example, between 0.60 and 0.70. These values can be combined into two independent stability coefficient records. Example: Observe the tasks processed by a certain approval personnel in the department from 8:00 to 18:00 during the day in the past 2 months and find that their on - time completion rate reaches 85%, then Set to 0.85, while the on-time completion rate is only 60% after 18:00 at night, then = 0.60, which is finally used for subsequent formula reference.
[0153] The acquisition steps of
[0154] This parameter represents the number of task records of the approver, which is obtained by counting in the completed tasks mentioned above. It is necessary to ensure that all records are approval tasks independently or mainly completed by this person before they can be included in the count of . If it is statistically found that this person has completed a total of 15 approval tasks with complete duration records in the past period of time, then record .
[0155] The acquisition steps of
[0156] This parameter represents the number of task categories of the approver. It is necessary to classify the task types handled by this person currently or historically. For example, the approval process is divided into multiple types such as contract approval, expense reimbursement, and procurement application. If this person is involved in a total of 4 different types of approval tasks, then , and for those who only handle a single type of approval within a period of time, may also occur. Acquisition example: Four types of approval services such as expense reimbursement, procurement application, legal review, and asset management are preset in the system. If a certain approver has handling records for all of them, then the number of his task categories is 4.
[0157] Calculation process:
[0158] In a specific scenario, historical and current task data of a certain approver have been obtained. The following are example values:
[0159] , the statistical results of the duration of 5 completed tasks of this person are hours, hours, hours, hours, hours.
[0160] , the current tasks of this person are divided into urgent and ordinary types, with a total of 2 types, corresponding to and .
[0161] (2 urgent cases), (3 regular cases), then .
[0162] , 4 unopened to-do tasks were found in the to-do list of this person.
[0163] hours, obtained by calculating the average completion time after extracting several completed tasks in the past three months, is 3 hours.
[0164] , after comparing with the department average efficiency and normalizing, the efficiency of this person is 0.80.
[0165] , the standard deviation of the processing time on different task types was statistically calculated and summarized to be 0.60.
[0166] 、 , the stability coefficient during weekdays and nights was measured.
[0167] First, calculate the numerator part:
[0168] ;
[0169] ;
[0170] Therefore, the numerator = = ;
[0171] Then, calculate the denominator part:
[0172] ;
[0173] ;
[0174] ;
[0175] Add these values together: ;
[0176] Then add 1: ;
[0177] So, the denominator = ;
[0178] Finally:
[0179] ;
[0180] When When it exceeds 2.00, it is marked as a high-load matching degree, indicating that the person is in a sensitive stage of work assignment under the comprehensive consideration of multiple dimensions such as the completed task duration, the current number of tasks, the number of tasks to be processed, and historical efficiency. If it falls below 1.00, it indicates that the person has a better balance between task load and efficiency.
[0181] This result shows that the person currently has a high load in multiple dimensions. If the approval task is continued to be assigned to this person, it may cause a certain degree of task backlog. When the value is getting larger and larger, it indicates that there is a more obvious contradiction between the current load and historical efficiency, and it needs to be carefully considered in the subsequent task assignment stage.
[0182] The steps to obtain the task processing ability index are as follows:
[0183] Based on the matching degree index of the approval personnel, set the task assignment threshold, screen the matching degree indexes of all approval personnel, extract the approval personnel with matching degree indexes higher than the task assignment threshold, and obtain the list of candidate approval personnel;
[0184] According to the list of candidate approval personnel, calculate the task completion duration of each candidate approval personnel. The calculation formula is:
[0185] ;
[0186] Among them, represents the task completion duration of the candidate approval personnel, is the total duration of the tasks completed by the candidate approval personnel in the past, represents the estimated duration of the current task to be processed, and represent the task switching time of the candidate approval personnel, is the standard processing duration of the candidate approval personnel's historical tasks, and represent the adjustment parameters of task complexity, 、 and represent the task processing efficiency of the candidate approval personnel in different time periods;
[0187] Based on the task completion duration, evaluate the task processing ability of the candidate approval personnel, and combine the task category, processing timeliness, and execution record to obtain the task processing ability index.
[0188] Specifically, based on the reviewer matching degree index, first specify an applicable task assignment threshold and load the matching degree indexes of all reviewers. After reading each matching degree index data one by one, refer to the previously collected execution samples of the approval process to determine a demarcation standard that can distinguish between high-load and medium-low-load personnel. If a specific threshold needs to be formed during the previous collection, statistical methods can be combined for segmented processing. For example, when selecting 50 past approval process records, task congestion is often more serious when the corresponding reviewer's final matching degree index ranges from 1.50 to 3.00. If the matching degree index is below 1.00, the load is generally relatively light. Therefore, 1.50 can be used as an empirical threshold to screen out personnel with higher matching degree indexes. Then, compare the matching degree indexes of all reviewers with this threshold, filter out those below this value, and retain the remaining records. For those personnel with special work patterns at different time periods, their average task processing efficiency and overtime hours will also be checked. If they meet the specific department requirements and their matching degree is higher than the set standard, these personnel will be listed as candidate objects. It is also necessary to further check their job qualifications or approval authority scope for those who meet the standards to confirm that they have the disposal ability for the corresponding task types. If their job responsibilities are incompatible, they will be excluded even if their matching degree indexes meet the standards. After the screening is completed, summarize the reviewers who pass the matching degree threshold to generate a list of candidate reviewers.
[0189] The benefit of the formula is that it comprehensively considers multiple factors such as the total duration of past completed tasks, the estimated duration of current tasks to be processed, task switching time, task complexity, and processing efficiency. Through the form of absolute difference, square root operation, and the superposition of multiple efficiencies at the denominator end, the impacts of different time periods and different task characteristics are incorporated into a unified measurement, providing more detailed data support for the subsequent evaluation of the overall task completion duration of candidate reviewers.
[0190] Steps to obtain the P parameter:
[0191] This parameter represents the total duration of past completed tasks of candidate reviewers. It is necessary to retrieve all the closed or archived tasks of this person in the system, and then extract the actual execution duration for each task record. Adding these duration values together can obtain the initial value of P. To avoid interference, abnormal records need to be excluded, such as time periods caused by overly long interruptions shown in the logs or tasks with missing records. Through multiple cross-checks, only effective processing durations are ensured to be included. For example, a certain candidate reviewer completed 8 tasks during the statistical period, which took 1.2 hours, 3.0 hours, 2.5 hours, 2.0 hours, 4.0 hours, 3.5 hours, 2.2 hours, and 2.6 hours respectively. Then, these numbers can be added together to get P = 1.2 + 3.0 + 2.5 + 2.0 + 4.0 + 3.5 + 2.2 + 2.6 = 21.0 hours.
[0192] Steps to obtain the Q parameter:
[0193] This parameter represents the estimated duration of the current task to be processed. It needs to be estimated based on information such as the operation process, workload, and existing historical processing data of the task before task allocation. For example, in an approval system, a typical processing duration is set for a certain process. The median value of similar tasks can be used to represent the average time required from receiving to completing the task, and then it can be appropriately increased or decreased in combination with the priority of the task or cross-departmental requirements. For example, a certain cross-departmental financial approval generally takes 2.5 hours on average, but the current task involves contract confirmation and requires additional query of supplier qualifications, which may increase by 1.0 hour on the basis of the basic time. So Q can be taken as 3.5 hours.
[0194] Steps to obtain the R parameter and the S parameter:
[0195] These two parameters represent the task switching time of candidate approvers. Usually, the monitor needs to measure the interval from the end of the previous task to the formal processing of the next task among different tasks, which can be calculated according to the timestamps recorded in the system operation log. Once the time difference between the previous task entering the archived state and the "start processing" flag of the next task is a switching duration. R and S can respectively reflect the average switching values in different types of tasks. For example, R refers to the average switching duration from normal approval to urgent approval, and S refers to the average switching duration from urgent approval to cross-departmental approval. These averages can be obtained by summarizing multiple switching processes within the statistical period and dividing by the number of switches. If the candidate approver has experience in multiple types of approvals and their common switching durations are within the range of 15 minutes to 30 minutes, these durations can be first converted into hours and then recorded as R and S respectively. For example, it is calculated that R = 0.3 hours and S = 0.4 hours.
[0196] Steps to obtain the W parameter:
[0197] This parameter is the standard processing duration of the candidate approver's historical tasks. It is necessary to count the actual processing times of the candidate approver in the past several same or similar types of tasks, extract these processing times respectively and calculate an arithmetic mean or median. Determine the processing method for outliers according to statistical principles. If it is found that some tasks are affected by external factors resulting in abnormally long or short processing times, they need to be excluded or corrected by means of mean regression as appropriate. After selecting the eligible task records, add up their processing durations and divide by the total number of records to obtain a basic average duration, and then fine-tune this value in combination with the timeliness evaluation given by the department or enterprise. For example, if a large number of statistics show that the median processing duration of this person for ordinary approval tasks is about 2.5 hours, then W can be set to 2.5, or it can be corrected through more samples.
[0198] Steps to obtain the X parameter and the Y parameter:
[0199] These two parameters represent the adjustment parameters for task complexity. It is necessary to break down the current task to be processed, list elements such as the number of main sub-links, document review requirements, or the number of cross-departmental collaborations, assign values to these elements in a quantitative manner, and then sum or normalize them to obtain X or Y. For example, X specifically targets the number of sub-links for document verification, and Y targets the number of cross-departmental or external collaboration links. The larger the sum of the two, the higher the overall difficulty of the task usually indicates. It can be compared with 50 similar tasks in the sample. If it is found that most tasks need to verify 2 to 3 documents and the department cooperation is 1 to 2 times, it can be regarded as the normal difficulty range. If this task needs to verify 5 documents and there are 4 cross-departmental communications, the sum of X and Y will increase significantly. When the statistical result is 2 or 3, it can be understood as medium complexity, and 5 or 6 is higher complexity.
[0200] Steps to obtain the Z parameter, A parameter, and B parameter:
[0201] These three parameters represent the task processing efficiency of candidate approvers at different time periods and are used for summing the denominator. It is necessary to distinguish and count the completion efficiency of personnel during the early shift, normal working hours, and evening or overtime periods. Divide the number of tasks completed within each time period by the total effective working hours of that time period to obtain a time-period efficiency value, and finally map it to the same interval and record it in Z, A, and B respectively. For example, Z represents the efficiency from 8:00 to 12:00, A represents the efficiency from 13:00 to 18:00, and B represents the efficiency after 19:00 or on weekends. If this person completes tasks quickly from 8:00 to 12:00 and can complete 0.4 pieces per hour within 2 weeks, it can be mapped to 0.75 after comparing with the department average efficiency. If 0.3 pieces can be completed per hour from 13:00 to 18:00, it is assigned 0.60, and if 0.2 pieces can be completed per hour during the evening period, it can be assigned 0.45.
[0202] Calculation process:
[0203] P = 21.0 hours (total time consumed for tasks completed by this candidate approver during the statistical period);
[0204] Q = 3.5 hours (estimated time required for the current task to be processed);
[0205] R = 0.3 hours, S = 0.4 hours (average duration for switching between different types of tasks);
[0206] W = 2.5 hours (standard processing duration obtained from historical records);
[0207] X = 2, Y = 1 (quantitative parameters obtained after breaking down task complexity);
[0208] Z = 0.75, A = 0.60, B = 0.45 (mapped values of efficiency for each time period);
[0209] First, calculate the numerator part:
[0210] ;
[0211] ;
[0212] ;
[0213] Add these three terms together: ;
[0214] Then, calculate the denominator part:
[0215] ;
[0216] So:
[0217] ;
[0218] This result indicates that after considering factors such as the total duration of previously completed tasks, the time required for current pending tasks, task switching time, and processing efficiency at different time periods, the overall task completion duration estimated for this candidate approver is approximately 10.38 hours. If a lower value is found for another candidate approver during subsequent comparisons, it means that person may be able to complete this task faster and may also be given priority in the assignment.
[0219] Based on the task completion duration, first separately check the historical records of each candidate approver and compare with the value calculated above. Arrange all candidate approvers in ascending order. At this time, it is necessary to conduct a supplementary review of the stability of the same person in different types of tasks. For example, if someone is extremely familiar with contract approval work, then when dealing with the same category, it may decrease significantly. If another person is good at the financial reimbursement process, then when facing financial-related approvals, it may be relatively low. Therefore, when making statistics, it is necessary to obtain the professional proficiency descriptions of their corresponding task categories and check the completion time period data of their corresponding task categories in historical approvals. If someone has a long processing duration or cross-departmental communication exceeding the predetermined standard for several consecutive times, it is necessary to track the time-consuming frequency in the communication link. If someone's average communication delay is within 0.5 hours each time, it is considered relatively stable, but if it exceeds 1.0 hour, it is marked as communication delay behavior. After confirming this information and combining it with the requirements of processing timeliness, if some approvals have urgent requirements within 48 hours, those Candidates with less than 10 hours are selected, and their overall performance in handling such tasks in their past execution records is preferably reviewed. Finally, all analyses are integrated to compare the task category suitability, past records, and corresponding time limit requirements of candidate approvers, so as to obtain task processing ability indicators.
[0220] The steps for obtaining the approval task assignment sequence are as follows:
[0221] Based on the task processing ability indicators, the task processing abilities of all candidate approvers are extracted, the task saturation of the candidate approvers is called, and the current task quantity and task load of each candidate approver are analyzed to obtain the task status information of the candidate approvers;
[0222] According to the task status information of the candidate approvers, a matching analysis is performed on the task processing ability indicators and task saturation of the candidate approvers, the acceptable task quantity of each candidate approver is analyzed, and the candidate approvers are ranked according to the acceptable task quantity to generate a candidate approver task assignment priority sequence;
[0223] Based on the candidate approver task assignment priority sequence, candidate approvers are sequentially selected according to the execution requirements of the approval tasks, the approval tasks are assigned in the order of priority, and the task assignment order is recorded to generate an approval task assignment sequence.
[0224] Specifically, based on the task processing ability indicators, first centrally read the historical and current work records of all candidate approvers, and check their average processing duration and recent processing speed for each task type according to the existing statistical data. Then, compare the task saturation of each candidate approver. The saturation generally needs to be summarized from aspects such as the number of ongoing tasks, the number of tasks to be done, and the proportion of urgent tasks for this approver in a recent period. For this, the current to-do list of each person can be queried in the database first and the tasks that have been assigned but not yet started can be marked. Then, combined with the data of the past several weeks or months, summarize the task completion situation in each time period. If the average number of ongoing tasks of this approver is about 5 in the selected reference period and at least 2 of them are in high-priority or cross-departmental collaboration scenarios, then define its saturation in a higher range. If specific high, medium, and low saturation levels need to be divided, the data of all approvers can be analyzed by quantiles with reference to statistical methods. For example, take 1.2 times the median of the past multiple handling records as the boundary. List the approvers who exceed this value as high saturation, list those who fall within a certain range below this value as medium saturation, and list those who are significantly lower than this value as low saturation. If an approver has been maintaining a high saturation level for a long time, it means that the workload he is currently undertaking is relatively large. After the saturation is statistically completed, then view the existing processing ability value of the corresponding approver side by side with its saturation. If an approver has a high ability value but the saturation is close to the critical point, it may not be appropriate to arrange more tasks. If another approver has a medium ability value and a low saturation, perhaps it is more suitable to undertake some new tasks in the current period. After aggregating the data of all candidate approvers by personnel number or name, bring them into the software to comprehensively sort out their current task quantity, the distribution of the time used for completed tasks, and the saturation index. Finally, obtain the task status information of the candidate approvers.
[0225] According to the candidate approver task status information, first extract the ability index values and saturation values of each candidate approver and conduct a matching analysis. Specifically, the ability index of a single approver can be combined with their current saturation through a one-by-one comparison method. If it is in the area with high ability index and low saturation, it can be initially determined that they have the potential to accept more tasks. If it is in the area with medium ability index and high saturation, it indicates that the approver has a relatively high work pressure recently and needs to be cautious when allocating additional tasks. To determine the number of tasks that each approver can accept, refer to the processing rate of the tasks they have completed in the past three weeks, divided by the hour dimension or day dimension. Compare the sum of the task processing rate and the remaining workload of the current in-progress tasks once. If the rate is significantly greater than the median value of the current remaining workload, it means that the number of tasks that can be accepted can be maintained at about 2 to 3. Otherwise, they may only be able to accept 1 more task or not accept new tasks. Such a judgment process can be cross-validated by combining several sample intervals. For example, first observe the processing rate curve of the approver in the most recent two weeks and the two weeks before that. If the difference between the two is less than 10%, it is confirmed that the evaluation conclusion of the number of tasks that can be accepted is relatively stable. If the difference exceeds 20%, an additional correction amount should be set in the calculation. After completing this judgment, record the number of tasks that each candidate approver can accept in sequence, and rank these approvers in descending order of the number of tasks that can be accepted. If there are multiple ties, the processing speed of these approvers in key category tasks can be compared again. Finally, generate the candidate approver task allocation priority sequence.
[0226] Based on the candidate approver task allocation priority sequence, first match the candidate approvers item by item according to the execution requirements of the approval tasks. Retrieve the list of approvers with the highest priority from the first position of the sequence and compare their available time slots and remaining energy for the same type of tasks in the current cycle. If the number of tasks they can accept is still available and matches the target task category, assign this task to this approver and record the allocation. Then, look at the next approver in the sequence and continue to match and allocate for the same type or similar type of approval projects. When encountering tasks that require cross-department co-signature or involve special data review, the completion duration records of the approver's previous handling of this type of task will be additionally called for confirmation. If it is found that their overtime frequency is too high or the completion duration is often higher than the benchmark time limit set by the department for the same type of tasks, give appropriate avoidance or lower the priority in the allocation. A time limit can also be set in advance to mark this benchmark time limit. For example, in most ordinary approval processes, the benchmark time limit is set at 24 hours, but in cross-department processes, it can be appropriately increased to 36 hours. Check whether the approver's previous completion duration for the same type is significantly deviated from this benchmark. If the deviation is greater than 6 hours for multiple times, automatically lower their sequence priority. Select the corresponding approvers for all the approval tasks yet to be allocated according to such a matching process. After all tasks are successfully allocated, record the assignment order corresponding to each task this time. Finally, generate the approval task allocation sequence.
[0227] The steps for obtaining the approval task assignment record are as follows:
[0228] Based on the approval task assignment sequence, extract the information of the first approver in the approval task assignment sequence, call the identity information and position information of the approver, and verify the approval task assignment sequence to generate the first assignment information of the approval task;
[0229] According to the first assignment information of the approval task, match the file identifier of the file to be approved, call the approval file content and assignment time data, bind the approval file to the first approver, record the assignment situation of the approval file, and generate the approval task assignment binding information;
[0230] Based on the approval task assignment binding information, establish an approval task assignment record, store the approval file assignment time, file identifier and the information of the first approver, and synchronously update the approval task management database.
[0231] Specifically, based on the approval task assignment sequence, first read the list of approved personnel arranged in order and locate the first approved personnel in the sequence. Then, compare the information such as the identity number, position number, and associated department code registered by this approved personnel in the system, and retrieve the verified record in the identity database in character form. If the matching degree of the identity information of this approved personnel is greater than the pre-established benchmark value, it is considered that the information is consistent. If the corresponding benchmark is not reached, it is marked as identity to be confirmed. This benchmark value is generally obtained by continuously observing the identity verification results of multiple approved personnel and counting the similarity. For example, when comparing the identity labels of 50 approved personnel with the system records, it is found that a matching degree of more than 90% can successfully identify. Therefore, 90% can be selected as this benchmark. Then, continue to retrieve the position name and authority scope corresponding to this approved personnel in the position description dataset, and combine the authority mapping table collected in the early stage to judge whether they have the processing authority involved in this approval process. If the authority code matches the process requirements, record it as an executable position. If it is found that there are restrictive clauses in the position description, such as only being able to handle financial tasks while the current task belongs to contract approval, mark it as position mismatch and abandon this assignment during subsequent execution. After completing the position confirmation, it is necessary to further verify the approval task assignment sequence, make a distinguishing mark between the number of this approved personnel and the remaining subsequent personnel in the sequence to prevent duplicate assignment. Then, query the available time period of this approved personnel on the current working day or work week according to the internal contact file. If there is a serious conflict with the task requirements, it will be prompted in the analysis result. This conflict threshold can be obtained from the scheduling tolerance range agreed within the department. For example, some departments stipulate that any overlap with the high-priority project duration reaching 80% or more is considered a serious conflict. The statistical method can be combined with the task schedule of this approved personnel. If the overlap degree reaches or exceeds 80%, it is listed as unavailable. If it is less than 80%, it can still be assigned. After all verifications are completed, organize and record the approved personnel number, identity description, and position information used in this assignment to obtain the first assignment information of the approval task.
[0232] According to the first-assigned information of the approval task, first retrieve the list of file identifiers in the current system that are in the pending approval status. Each file has a unique identifier number and its basic attributes in the archival record, such as file type, proposed handler information, etc. Then, use the file identifier number as the primary key to query the corresponding approval file content in the database, and confirm whether it contains fields such as contract attachments or amount terms. If it is detected that the correspondence between these fields and the position permissions is consistent, the further process can be entered. If the position permissions do not support processing such fields, it should be marked as non-bindable in the subsequent records. After the matching is completed, extract the allocation time data of the current system. This allocation time can be combined with the allocation timestamp of the superior dispatching center to form an accurate task assignment moment. If it is found that the stored timestamp deviates significantly from the regular time allocation interval (for example, more than 2 hours, and the acceptable interval set within the department is usually in the range of 0 to 2 hours, and this value is obtained by statistically averaging the actual time consumption during the past dozens of task issuances plus the standard deviation), then mark this record as an abnormal time point. If the difference from the acceptable interval stipulated by the department does not exceed 2 hours, it is regarded as a normal allocation operation. Subsequently, associate the approval file with the first-assigned information of the current approval task one by one and summarize the key information during the binding process, including the approval personnel number, position name, file identifier, and specific allocation moment. After the summarization, update the corresponding database record to obtain the approval task allocation binding information.
[0233] Based on the approval task allocation binding information, first collect multiple data items including the approval file allocation time, file identifier, and the number of the first approval personnel, and order them in a centralized approval process tracking form. Classify and mark these binding information according to different stages such as in progress and completed. If the file has entered a certain collaborative stage in the previous link, its temporary status will also be translated as "in progress across departments". If it is in this status currently, it will not be updated repeatedly temporarily to prevent overwriting existing process traces. Subsequently, use the system serial number as the retrieval clue to associate all newly created binding information with the approval personnel identity ID, and write these combined data into the approval task management database. Each piece of data carries its effective moment to indicate the precise recording point. The recording method can design an extended field in the database to store this moment. If it is necessary to summarize all allocation behaviors within a certain time period during statistics, the corresponding list can be directly obtained through moment query. Here, compare the previously stored allocation time with the current server time. If the time difference is within the range of 0 to 24 hours, it is listed as a newly created allocation on the same day. If it exceeds 24 hours, it is classified as a historical allocation. After grouping the allocation situations of various categories, it can also be provided for the management to review. When the above grouping and sorting are all completed, update the database index list and end this round of operation to obtain the final approval task allocation record.
[0234] The above are only the preferred embodiments of the present invention, and do not limit the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A document approval method based on NLP, characterized in that: The following steps are involved: Input the approval document, perform word segmentation on the content of the approval document, calculate the similarity between the word segmentation result and the preset business type dictionary, and generate the file business classification result; based on the file business classification result, perform semantic vector conversion on the keywords in the approval document content, calculate the semantic association strength between the keywords, and generate the file semantic feature vector; Based on the document semantic feature vector, the responsibilities of each department and the responsibilities of each person in the organizational structure are called, the matching values of the vector content and the responsibilities of the department and the person are calculated, and the responsibility matching results are generated; Based on the responsibility matching results, the current number of tasks and task processing time of each approval personnel are counted, and the approval personnel matching index is calculated and generated; Based on the approval personnel matching index, a task allocation threshold is set, candidate approvers are screened according to the task allocation threshold, and a task processing capability index is generated; based on the task processing capability index and combined with the task saturation of the candidate approvers, an approval task allocation sequence is generated; Based on the approval task allocation sequence, extract the information of the first approver in the sequence, allocate the approval document to the first approver, record the allocation time and document identifier, and generate an approval task allocation record; The steps for obtaining the file semantic feature vector are as follows: Based on the file business classification result, the approval file content is parsed, keywords in the approval file content are extracted, keywords are screened, redundant keywords, synonyms or interference items are removed, and keywords are normalized to obtain a keyword list; According to the keyword list, the semantic association strength between the keywords is calculated, and the calculation formula is: ; in, Represents the strength of semantic association between keywords. Representative keywords The frequency of occurrence in the text, and Respectively represent and The semantic vector value of keywords, and Respectively represent and The relative position deviation of keywords in the context window, and Represents the semantic dependency distance of keywords in the text, and Represents the dependency weight of keywords in different syntactic structures, and Represents the number of times a keyword appears in different text fragments; Based on the semantic association strength, the semantic association strengths corresponding to all keyword combinations are collected, and the distribution statistics of the semantic association strengths in the range of 0 to 1 are performed. If the semantic association strength of some keyword pairs exceeds 0.70, they can be classified as a high association group. If it is between 0.30 and 0.70, it is marked as a moderate association range. If it is lower than 0.30, it is considered to be weakly associated. Keywords belonging to the high association group are aggregated and recorded in the same keyword cluster. Then, the main type labels in the business classification results are compared one by one to see if there is a keyword cluster corresponding to a specific business category. If it is determined that there is an overlap, the keyword cluster is bound to the category. Finally, all keywords and cluster affiliations are summarized to construct a file semantic feature vector.
2. The NLP-based document approval method according to claim 1, characterized in that: The steps for obtaining the file service classification result are: After the approval document is input, it is segmented through natural language processing, the text in the document is semantically analyzed, and the sentences are separated into independent vocabulary units to form the segmentation results; According to the word segmentation result, the similarity between each word segmentation element and an entry in a preset business type dictionary is calculated to obtain a similarity result; Based on the similarity result, the similarity scores of each business type are compared, and the one with the highest similarity score is selected as the type identifier of the file to generate a file business classification result.
3. The NLP-based document approval method according to claim 1, characterized in that: The steps for obtaining the responsibility matching result are: Based on the document semantic feature vector, call the responsibilities of each department and the responsibilities of each person in the organizational structure, parse the responsibilities description text, extract the responsibilities keywords, construct the responsibilities keyword vector, and obtain the responsibilities keyword vector matrix; According to the responsibility keyword vector matrix, the matching value between the document semantic feature vector and the responsibility keyword vector is calculated, and the expression is: ; in, Represents the matching value, is the first part of the document semantic feature vector Dimensional components, is the first Dimensional components, and Represent the relative distances between document keywords and responsibility keywords in the semantic space, For the The frequency value of the matching words, and Represents the position index of the document keywords and responsibility keywords in the syntactic structure, , and Represents the matching similarity coefficient between each job description text, Represents the number of dimensions of the document semantic feature vector, Represents the number of matching words in the responsibility keyword vector, The total number of job description texts; Based on the matching value, a matching evaluation is performed on the job descriptions of each department and the job descriptions of personnel positions, and the responsibility item with the highest matching value is selected as the belonging department and responsible person to obtain a responsibility matching result.
4. The NLP-based document approval method according to claim 1, characterized in that: The steps for obtaining the approval personnel matching index are as follows: Based on the responsibility matching results, the current number of tasks of each approver is counted, and the cumulative task processing time of each approver's processed tasks is calculated. At the same time, the list of pending tasks of the approver is extracted to obtain the task statistics information of the approver; According to the task statistics of the reviewers, the task processing mode of each reviewer is analyzed, the average number of tasks completed per unit time is calculated, and the current task saturation degree of each reviewer is counted to obtain the reviewer task load index; Based on the task load index of the approval personnel, the matching index of the approval personnel is calculated. The calculation formula is: ; in, Represents the approval personnel matching index, The task processing time of the approved personnel who have completed the task. The current number of tasks for the approver. Represents the number of tasks to be processed by the approver. Represents the average completion time of historical tasks of the approver. Represents the task processing efficiency of the approver on different task types, Represents the standard deviation of task processing time of approvers in different task categories. and Represents the stability coefficient of task completion of the approver in different time periods, Represents the number of task records of the approver. Represents the number of task categories for approvers.
5. The NLP-based document approval method according to claim 1, characterized in that: The steps for obtaining the task processing capability indicator are as follows: Based on the matching index of the approver, a task allocation threshold is set, the matching indexes of all approvers are screened, and the approvers whose matching indexes are higher than the task allocation threshold are extracted to obtain a list of candidate approvers; According to the list of candidate reviewers, the task completion time of each candidate reviewer is calculated, and the calculation formula is: ; in, Represents the task completion time of the candidate approver, is the total time that candidate approvers have completed tasks in the past. Represents the estimated duration of the current pending task. and Indicates the task switching time of candidate approvers. The standard processing time of the candidate approver's historical tasks. and A tuning parameter representing the complexity of the task, , and Represents the task processing efficiency of candidate approvers in different time periods; Based on the task completion time, the task processing capability of the candidate approver is evaluated, and the task processing capability index is obtained by combining the task category, processing timeliness and execution record.
6. The NLP-based document approval method according to claim 1, characterized in that: The steps for obtaining the approval task allocation sequence are as follows: Based on the task processing capability indicator, the task processing capabilities of all candidate approvers are extracted, the task saturation of the candidate approvers is called, and the current number of tasks and task load of each candidate approver are analyzed to obtain the task status information of the candidate approver; According to the task status information of the candidate approvers, matching analysis is performed on the task processing capability index and task saturation of the candidate approvers, the acceptable number of tasks for each candidate approver is analyzed, and the candidate approvers are prioritized based on the acceptable number of tasks to generate a task allocation priority sequence for the candidate approvers; Based on the candidate approver task allocation priority sequence, candidate approvers are selected in turn according to the execution requirements of the approval tasks, the approval tasks are allocated in order of priority, and the task allocation order is recorded to generate an approval task allocation sequence.
7. The NLP-based document approval method according to claim 1, characterized in that: The steps for obtaining the approval task allocation record are as follows: Based on the approval task allocation sequence, extract the information of the first approver in the approval task allocation sequence, call the identity information and position information of the approver, and verify the approval task allocation sequence to generate the first allocation information of the approval task; According to the approval task first allocation information, the file identifier of the to-be-approved file is matched, the approval file content and allocation time data are called, and the approval file is bound to the first approver, the allocation of the approval file is recorded, and the approval task allocation binding information is generated; Based on the approval task allocation binding information, an approval task allocation record is established to store the approval document allocation time, document identification and first approver information, and the approval task management database is updated synchronously.
Citation Information
Patent Citations
File approval method and device, medium and calculation equipment
CN108376333A
Extracting actionable items from documents and assigning the actionable items to responsible parties
US20210034703A1