Work order feature classification method and system based on natural language processing
By obtaining the salesperson's historical and current job identification, identifying common ambiguous feature sets and utilizing natural language processing algorithms, the problem of ambiguous feature recognition in the work order classification system during job changes is solved, and accurate classification and efficient processing of work order features are achieved.
Patent Information
- Application Number
- CN202511202700.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-08-27
AI Technical Summary
The existing work order classification system has difficulty accurately identifying and distinguishing the mixing of cross-position terminology after business personnel change positions, resulting in incorrect classification of work order features.
By obtaining the salesperson's historical and current job identification, we identify common ambiguous feature sets, use natural language processing algorithms to perform text recognition on work order texts, and combine adjacent features to comprehensively determine the meaning of the text for classification.
Improves the accuracy and efficiency of work order feature recognition, ensures the accuracy and efficiency of work order classification after enterprise architecture updates, promptly identifies job changes, and focuses on key ambiguous features.
Smart Images

Figure CN120705316A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of language processing, and in particular to a work order feature classification method and system based on natural language processing. Background Art
[0002] Work orders carry important information such as customer requests and business processing instructions. Against the backdrop of enterprise architecture updates, the shift in collaboration models between business departments and the reorganization of business processes have placed higher demands on the accurate recording and classification of work orders to ensure that information can be efficiently and accurately transmitted between different positions and functions. Currently, enterprises generally adopt work order classification systems based on natural language processing technology. These systems use text analysis algorithms to extract and identify features such as keywords and semantic structures in the work order text content, thereby achieving automatic classification of work order features. At the same time, when the enterprise architecture is updated, the job information of business personnel is also updated more promptly. The system can perform preliminary business classification of work order texts based on the current job identification of business personnel, such as classifying work orders from the sales department into the sales category and work orders from the after-sales department into the after-sales category. This makes the work order classification consistent with the business process to a certain extent, providing certain convenience for subsequent business processing.
[0003] With the continuous updating of enterprise architecture and rapid business development, business personnel frequently change positions. As a result of these changes, while the work order classification system can classify work orders based on job information, when recording work orders, business personnel may subconsciously use the professional terms of their previous position to describe the work content of their current position. Existing systems often struggle to accurately identify and distinguish this mixed use of cross-position terminology. This results in work order characteristics of the current position being incorrectly classified into the characteristic categories of the previous position, leading to the problem of misclassification of work orders.
[0004] Prior art (CN114528399A) discloses a work order text classification method, device, storage medium, and computer equipment. However, when recording work orders after a business person changes positions, this prior art uses the professional terminology of the previous position to describe the current position's work content. This method cannot accurately distinguish between mixed terminology, resulting in low accuracy in work order feature recognition.
[0005] Therefore, how to improve the accuracy of work order feature recognition is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention
[0006] In order to address the deficiencies in the prior art, the present application provides a work order feature classification method and system based on natural language processing to solve the technical problem of inaccurate work order feature classification.
[0007] This application adopts the following technical solution.
[0008] In a first aspect, the present application provides a work order feature classification method based on natural language processing, comprising: In response to a work order processing instruction, determining a salesperson identifier corresponding to the target work order based on target work order data, wherein the work order processing instruction indicates starting to process the target work order data included in the target work order; Obtaining a historical position identifier and a current position identifier corresponding to the salesperson identifier; when the historical position identifier and the current position identifier are different, obtaining a common ambiguous feature set corresponding to both the current position identifier and the historical position identifier, wherein the common ambiguous feature set includes a plurality of common ambiguous features with the same text characters but different text meanings, wherein the different text meanings indicate that the common ambiguous features have different text meanings when corresponding to different positions; Performing text recognition on the target work order data using a natural language processing algorithm to obtain a plurality of text features; and for each of the text features, if the text feature belongs to the common ambiguous feature set, determining at least one adjacent feature corresponding to the text feature; An available text meaning corresponding to the text feature is determined according to the text feature and the at least one adjacent feature, so as to classify the text feature in the target work order based on the available text meaning.
[0009] Optionally, obtaining a common ambiguous feature set corresponding to both the current job identifier and the historical job identifier includes: Determine a current feature identifier set corresponding to the current position identifier and a historical feature identifier set corresponding to the historical position identifier, wherein both the current feature identifier set and the historical feature identifier set include correspondences between a plurality of text characters and text meanings; Determining, based on the current feature identification set and the historical feature identification set, initial ambiguous features of a plurality of text characters having the same characters, and obtaining a plurality of text meanings corresponding to each of the initial ambiguous features; For each of the initial ambiguous features, when the meanings of the multiple texts corresponding to the initial ambiguous feature are different, the initial ambiguous feature is determined to be a common ambiguous feature.
[0010] Optionally, determining at least one adjacent feature corresponding to the text feature includes: Determining, based on the plurality of text features, a natural sentence corresponding to the text feature, wherein the natural sentence includes a plurality of text words, and the text word is the text feature or any adjacent entity; The text words are combined in pairs to obtain at least one entity discrimination group, and the adjacent features are obtained based on the entity discrimination group having actual meaning.
[0011] Optionally, determining at least one adjacent feature corresponding to the text feature includes: Determining, based on the plurality of text features, a natural sentence corresponding to the text feature, wherein the natural sentence includes a plurality of text words, and the text word is the text feature or any adjacent entity; Randomly combine the text words in a quantity M to obtain entity discrimination groups corresponding to the quantity M; the quantity M=3,…,N, where N is a positive integer; According to each entity discriminant group corresponding to the quantity M or the quantity 2, the entity discriminant group having actual meaning is used as the adjacent feature.
[0012] Optionally, determining at least one adjacent feature corresponding to the text feature includes: The search interval for adjacent features is defined, with the location of the text feature as the center and expanding to the preset range before and after it. Within the search interval, candidate adjacent features that are semantically related to the text feature are preliminarily screened based on the text's grammatical structure, semantic representation, and logical relationships in the business scenario. Using the semantic analysis model, the semantic relevance between the candidate adjacent features and the text feature is quantitatively evaluated, and the co-occurrence probability between the candidate adjacent features and the text feature is calculated. When the co-occurrence probability is not less than the preset semantic relevance strength threshold, it is determined to be an adjacent feature.
[0013] Optionally, the preliminarily screening candidate adjacent features that are semantically associated with the text feature within the search interval based on the grammatical structure, semantic representation, and logical relationship of the text in the business scenario includes: Perform feature extraction on the text feature and the candidate feature text in the search interval to obtain grammatical features and semantic features of the text, wherein the grammatical features are TF-IDF feature vectors, and the semantic features are semantic vector representations of the text obtained using a pre-trained language model; Calculate the similarity between the text feature and the candidate feature text in terms of grammatical structure, semantic representation, and logical relationship in the business scenario, and obtain grammatical structure similarity, semantic representation similarity, and business logic similarity; Perform weighted fusion on the grammatical structure similarity, semantic representation similarity, business logic similarity and their corresponding weights to obtain the comprehensive similarity score of the candidate feature text; When the comprehensive similarity score is greater than a preset score threshold, the candidate feature text corresponding to the comprehensive similarity score is confirmed as a candidate adjacent feature.
[0014] Optionally, before determining the current feature identifier set corresponding to the current position identifier, the method further includes: Acquire an initial feature identifier set, wherein the initial feature identifier set includes a historical current feature identifier set; Determining whether the plurality of text features are new text features compared to the initial feature identification set; If so, the initial feature identification set is updated according to the new text features to obtain the current feature identification set.
[0015] Optionally, the natural language processing algorithm runs based on a word vector model algorithm.
[0016] In a second aspect, the present application provides a work order feature classification system based on natural language processing, comprising: A first identification module is configured to determine, in response to a work order processing instruction, a salesperson identifier corresponding to a target work order based on target work order data, wherein the work order processing instruction indicates starting to process the target work order data included in the target work order; an acquisition module, configured to acquire a historical position identifier and a current position identifier corresponding to the salesperson identifier; and when the historical position identifier and the current position identifier are different, acquire a common ambiguous feature set corresponding to both the current position identifier and the historical position identifier, wherein the common ambiguous feature set includes a plurality of common ambiguous features having the same text characters but different text meanings, wherein the different text meanings indicate that the common ambiguous features have different text meanings when corresponding to different positions; a second recognition module configured to perform text recognition on the target work order data using a natural language processing algorithm to obtain a plurality of text features; and for each of the text features, if the text feature belongs to the common ambiguous feature set, determine at least one adjacent feature corresponding to the text feature; A processing module is configured to determine, based on the text feature and the at least one adjacent feature, an available text meaning corresponding to the text feature, so as to classify the text feature in the target work order based on the available text meaning.
[0017] Optionally, the acquisition module, when acquiring the common ambiguous feature set corresponding to both the current job identifier and the historical job identifier, is configured to: Determine a current feature identifier set corresponding to the current position identifier and a historical feature identifier set corresponding to the historical position identifier, wherein both the current feature identifier set and the historical feature identifier set include correspondences between a plurality of text characters and text meanings; Determining, based on the current feature identification set and the historical feature identification set, initial ambiguous features of a plurality of text characters having the same characters, and obtaining a plurality of text meanings corresponding to each of the initial ambiguous features; For each of the initial ambiguous features, when the meanings of the multiple texts corresponding to the initial ambiguous feature are different, the initial ambiguous feature is determined to be a common ambiguous feature.
[0018] Optionally, the second recognition module, when determining at least one adjacent feature corresponding to the text feature, is configured to: Determining, based on the plurality of text features, a natural sentence corresponding to the text feature, wherein the natural sentence includes a plurality of text words, and the text word is the text feature or any adjacent entity; The text words are combined in pairs to obtain at least one entity discrimination group, and the adjacent features are obtained based on the entity discrimination group having actual meaning.
[0019] Optionally, the above system further includes: Update modules for: Acquire an initial feature identifier set, wherein the initial feature identifier set includes a historical current feature identifier set; Determining whether the plurality of text features are new text features compared to the initial feature identification set; If so, the initial feature identification set is updated according to the new text features to obtain the current feature identification set.
[0020] Optionally, the second recognition module, when determining at least one adjacent feature corresponding to the text feature, is configured to: Determining, based on the plurality of text features, a natural sentence corresponding to the text feature, wherein the natural sentence includes a plurality of text words, and the text word is the text feature or any adjacent entity; Randomly combine the text words in a quantity M to obtain entity discrimination groups corresponding to the quantity M; the quantity M=3,…,N, where N is a positive integer; According to each entity discriminant group corresponding to the quantity M or the quantity 2, the entity discriminant group having actual meaning is used as the adjacent feature.
[0021] Optionally, the natural language processing algorithm runs based on a word vector model algorithm.
[0022] In a third aspect, the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the work order feature classification method based on natural language processing as described in any one of the first aspects.
[0023] In a fourth aspect, the present application provides a computer-readable storage medium, which stores computer execution instructions. When the computer execution instructions are executed by a processor, they are used to implement the work order feature classification method based on natural language processing as described in any one of the first aspects.
[0024] Compared with the prior art, the beneficial effects of this application include at least: The work order feature classification method based on natural language processing provided in this application can accurately identify the job changes of salesmen by judging whether the historical job identification and the current job identification are the same, and accordingly determine the common ambiguous feature set corresponding to the current position and the historical position; after the work order text is recognized by the natural language processing algorithm, the text features belonging to the ambiguous feature set are combined with their adjacent features to comprehensively determine the meaning of the text, and then classify it according to the meaning. This process can not only detect job changes in a timely manner and focus on key ambiguous features, but also deeply understand the text based on the context and accurately classify it. At the same time, it can simultaneously recognize other texts when determining the salesman, improve the efficiency of text analysis, and effectively solve the problem of incorrect work order classification caused by the use of old terminology due to job changes of sales personnel after the enterprise structure is updated, and ensure the accuracy and efficiency of work order processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0026] Figure 1 A flowchart of a work order feature classification method based on natural language processing provided in an embodiment of the present application; Figure 2 A schematic diagram of the structure of a work order feature classification system based on natural language processing provided in an embodiment of the present application; Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0027] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0028] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of systems and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0029] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0030] like Figure 1 As shown, Figure 1 This is a flow chart of a method for classifying work order features based on natural language processing provided in an embodiment of the present application. The execution subject of the embodiment of the present application may be an electronic device. Example 1 of the present application provides a method for classifying work order features based on natural language processing, which may specifically include steps S201 to S204, wherein: S201 . In response to a work order processing instruction, determine a salesperson identifier corresponding to a target work order based on target work order data, wherein the work order processing instruction indicates starting to process the target work order data included in the target work order.
[0031] Work order processing instructions should clearly include key information such as the target work order's unique identifier, salesperson identification information, and the storage location of the work order data. The salesperson identification can be identified from the preset header position of the target work order data packet. The header is designed with a fixed-length field, and the byte order of the identifier field strictly adheres to the network byte order specification. Identification can be achieved by parsing the header field of the target work order data packet and, based on the preset identifier encoding rules, decoding the byte sequence of the salesperson identification field in the header and converting it into a salesperson identification format that can be recognized within the system to ensure accurate extraction of salesperson information.
[0032] S202. Obtain a historical position identifier and a current position identifier corresponding to the salesperson identifier determined in S201; when the historical position identifier and the current position identifier are different, obtain a common ambiguous feature set corresponding to both the current position identifier and the historical position identifier, wherein the common ambiguous feature set includes multiple common ambiguous features with the same text characters but different text meanings, and the different text meanings indicate that the text meanings of the common ambiguous features are different when they correspond to different positions.
[0033] To obtain historical and current position identification, it is necessary to access the salesperson's electronic file in the enterprise's human resources management system or business operation record database and extract the identification information. The electronic file should include the salesperson's complete career history to ensure accurate access to their historical and current position identification.
[0034] Furthermore, text characters refer to the specific characters that make up text vocabulary, including basic units such as letters, numbers, Chinese characters, and special symbols. Text meaning refers to the meaning of the vocabulary composed of text characters in a specific context. For example, "IT" in a technical R&D position refers to "information technology," involving software development, etc.; in a customer service position, it refers to "information technology support," focusing on resolving customer technical issues. Another example is "legal affairs," in a legal and compliance department, refers to "legal affairs processing," covering contract review, etc.; in a sales department, it refers to "legal support," primarily involving sales compliance review of contract terms.
[0035] In one possible implementation, obtaining the common ambiguous feature set corresponding to both the current job identifier and the historical job identifier includes: Determine a current feature identifier set corresponding to the current position identifier and a historical feature identifier set corresponding to the historical position identifier, wherein both the current feature identifier set and the historical feature identifier set include correspondences between a plurality of text characters and text meanings; Determining, based on the current feature identification set and the historical feature identification set, initial ambiguous features of a plurality of text characters having the same characters, and obtaining a plurality of text meanings corresponding to each of the initial ambiguous features; For each of the initial ambiguous features, when the meanings of the multiple texts corresponding to the initial ambiguous feature are different, the initial ambiguous feature is determined to be a common ambiguous feature.
[0036] For example, in a certain company, suppose a business person previously worked in the "Software Development" position and later transferred to the "IT Support" position. In the current feature identifier set corresponding to their current "IT Support" position, "code" refers to "program instructions" and "debugging" refers to "troubleshooting equipment or system failures." In the historical feature identifier set corresponding to their historical "Software Development" position, "code" refers to "instructions in the software programming language" and "debugging" refers to "testing and correcting software." Comparing these two feature identifier sets reveals that the text characters "code" and "debugging" exist in both positions, but their meanings are different. Therefore, they are identified as common ambiguous features.
[0037] In one possible implementation, before determining the current feature identifier set corresponding to the current position identifier, the method further includes: Acquire an initial feature identifier set, wherein the initial feature identifier set includes a historical current feature identifier set; Determining whether the plurality of text features are new text features compared to the initial feature identification set; If so, the initial feature identification set is updated according to the new text features to obtain the current feature identification set.
[0038] The work order feature classification method based on natural language processing provided in this embodiment updates the feature identification set when new undefined text appears. This allows the system to incorporate new text features into the feature identification set without manual annotation, allowing it to be used to subsequently determine the common ambiguous feature set, thereby improving the accuracy and efficiency of work order feature classification.
[0039] S203. Perform text recognition on the target work order data using a natural language processing algorithm to obtain multiple text features; and for each text feature, if the text feature belongs to the common ambiguous feature set obtained in S202, determine at least one adjacent feature corresponding to the text feature.
[0040] The natural language processing algorithm is based on the word vector model algorithm. The target work order data is input into the word vector model for recognition to obtain multiple text features.
[0041] It should be noted that in this embodiment, the process of identifying the salesperson identification in S201 and the process of identifying the text features in S203 can be performed simultaneously, which can ensure the efficiency of text analysis on the one hand and improve the accuracy of text analysis on the other hand.
[0042] It can be understood that based on the linear sequence characteristics of text, the search interval of adjacent features is defined by taking the position of the text feature as the center and expanding to the previous and next text contents within a reasonable range.
[0043] Normally, the forward search can be set to the first A text features, and the backward search can be set to the last B text features, where A and B are positive integers determined in advance based on the semantic coherence of the text and experimental experience. This will frame the range in which adjacent features may appear, ensuring that it covers content that has a close semantic relationship with the text feature, while avoiding excessively expanding the search range and introducing irrelevant information.
[0044] Within the defined search interval, the initially identified adjacent text features are reviewed one by one. Based on multi-dimensional criteria such as the text's grammatical structure, semantic representation, and logical relationships within the business scenario, candidate adjacent features that may be semantically related to the text feature are initially screened. Features with no clear or very weak semantic connection to the text feature are eliminated, resulting in a more focused candidate list of adjacent features.
[0045] In one possible implementation, based on the above multi-dimensional criteria such as the grammatical structure, semantic representation, and logical relationships in the business scenario of the text, candidate adjacent features that may have semantic associations with the text feature are initially screened, including: Feature extraction is performed on the target text to be analyzed and the candidate feature text respectively to obtain the grammatical features and semantic features of the text.
[0046] Among them, the grammatical feature is the TF-IDF (Term Frequency-Inverse Document Frequency) feature vector, which can be obtained by segmenting the text, tagging the part of speech, extracting core words and calculating them. The semantic feature can use a pre-trained language model to obtain the semantic vector representation of the text.
[0047] The similarity between the target text and the candidate feature text is calculated in three dimensions: grammatical structure, semantic representation, and business logic.
[0048] Among them, the grammatical structure similarity is the core word overlap and the TF-IDF vector cosine similarity, and the ratio of the number of core words shared by the target text and the candidate feature text to the total number of core words in the target text and the candidate feature text is calculated as the core word overlap; the semantic representation similarity is the cosine similarity of the semantic vector, and the business logic similarity is the correlation between business keywords and business processes. The correlation can be the frequency of overlap between business keywords and preset keywords included in the business process.
[0049] According to the preset weight parameters, the similarity scores of the three dimensions are weighted and fused to obtain the comprehensive similarity score between the texts; the comprehensive similarity score is compared with the preset score threshold, and the candidate features with scores higher than the preset score threshold are screened out as adjacent features that may have semantic associations with the target text.
[0050] Furthermore, for candidate adjacent features, semantic analysis models in natural language processing algorithms, such as word vector models and pre-trained language models, are used to quantitatively evaluate the semantic relevance between candidate adjacent features and the text feature. By calculating the probability index of co-occurrence between the two, it is verified whether it meets the preset semantic relevance strength threshold. Specifically, the above-mentioned co-occurrence probability can be obtained by inputting the candidate adjacent features and the text feature into a co-occurrence statistical model. After obtaining the co-occurrence probability, the above-mentioned co-occurrence probability is compared with the preset semantic relevance strength threshold. If it is not less than the above-mentioned preset semantic relevance strength threshold, it is determined to meet the threshold requirements. For candidate adjacent features that meet the threshold requirements, they are formally determined as adjacent features of the text feature; those that do not meet the requirements are excluded, so as to ensure that the determined adjacent features can truly reflect the semantic relevance of the text feature in the context, and provide a reliable basis for the subsequent accurate determination of the available text meaning of the text feature.
[0051] In one possible case, during text processing, it is usually necessary to mark the positions of words in order to identify adjacent features. Generally speaking, each word is assigned a position identifier, and the words corresponding to adjacent position identifiers are considered adjacent. However, there is a special case in actual applications, that is, adjacent entities are not equivalent to adjacent features. For example, in the text "a box of apples and pears", the adjacent entities include "apples", "pears", "apple pears", and "a box", while the actual adjacent features should be "apples and pears" and "a box". At this time, if we still make a simple judgment based on the position identifier, it is easy to misjudge the adjacent features, resulting in deviations in the semantic understanding of the text and affecting the accuracy of subsequent text analysis. Therefore, there is an urgent need for more accurate methods to determine adjacent features to avoid such misjudgments.
[0052] In one implementable manner, determining at least one adjacent feature corresponding to the text feature includes: Determining, based on the plurality of text features, a natural sentence corresponding to the text feature, wherein the natural sentence includes a plurality of text words, and the text word is the text feature or any adjacent entity; The text words are combined in pairs to obtain at least one entity discrimination group, and the adjacent features are obtained based on the entity discrimination group having actual meaning.
[0053] For example, consider the natural sentence "It's raining today, so I performed maintenance on device A." The text features include "today," "weather," "rain," "I," "to," "device A," and "operation and maintenance." Based on these text features, the natural sentence corresponding to this text feature is determined to be "It's raining today, so I performed maintenance on device A." A natural sentence contains multiple text terms. These terms can be text features themselves, such as "today" and "weather," or adjacent entities, such as "device A rain," and "operation and maintenance." These text terms are then paired together to form multiple entity discriminant groups, such as "today's weather," "it's raining," "it's raining, I," "I performed maintenance on device A," and "performed maintenance on device A." These entity discriminant groups are then filtered based on whether they have practical meaning. For example, "it's raining" represents the weather condition, while "performed maintenance on device A" represents the action object and action type. These combinations have practical meaning and are therefore identified as adjacent features. Combinations like "it's raining, I" lack practical meaning and are therefore excluded. Finally, we obtain adjacent features such as "rainy weather" and "operation and maintenance of equipment A". These adjacent features can more accurately reflect the semantic information of the text and facilitate subsequent text analysis and processing.
[0054] In another possible implementation, determining at least one adjacent feature corresponding to the text feature includes: Determining, based on the plurality of text features, a natural sentence corresponding to the text feature, wherein the natural sentence includes a plurality of text words, and the text word is the text feature or any adjacent entity; Randomly combine the text words in a quantity M to obtain entity discrimination groups corresponding to the quantity M; the quantity M=3,…,N, where N is a positive integer; According to each entity discriminant group corresponding to the quantity M or the quantity 2, the entity discriminant group having actual meaning is used as the adjacent feature.
[0055] The work order feature classification method based on natural language processing provided by the embodiment of the present application, first of all, the pairwise combination method can directly and quickly exhaust the possibilities of adjacent combinations between text words, which is especially suitable for short texts or scenarios with high real-time requirements, and can give results in a relatively short time, which is suitable for preliminary text feature mining; the pairwise combination of adjacent words can often directly reflect the close semantic association between words, which helps to identify common phrase structures, has a natural advantage in identifying fixed collocations, common phrases, etc., and can effectively capture the basic semantic units in the text. The random number combination method can flexibly adapt to the needs of text analysis of different lengths and complexities. When M takes different values, semantic information at different levels can be mined, from the phrase level to the paragraph level, with stronger adaptability and scalability, and can capture text features more comprehensively.
[0056] S204 : Determine an available text meaning corresponding to the text feature according to the text feature and the at least one adjacent feature obtained in S203 , so as to classify the text feature in the target work order based on the available text meaning.
[0057] It's important to note that the above technical solution determines whether a salesperson has recently changed positions. If so, it obtains the salesperson's historical and current position information and identifies a set of common ambiguous words shared by both the historical and current positions. Subsequently, a vocabulary recognition process is executed on the current work order text, comparing each word in the text with the aforementioned ambiguous vocabulary set. If a word in the text falls within this ambiguous vocabulary set, the system further identifies at least one adjacent word in the text. By comprehensively considering the ambiguous word and its adjacent words, the system accurately determines the characteristic vocabulary category of the position to which the ambiguous word belongs.
[0058] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required by this application.
[0059] It should be further noted that, although the various steps in the flowchart are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps may be performed in other orders. Moreover, at least a portion of the steps in the flowchart may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but may be performed at different times. The execution order of these sub-steps or stages is not necessarily to be performed in sequence, but may be performed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0060] Figure 2 A structural diagram of a work order feature classification system based on natural language processing provided in an embodiment of the present application is shown as follows: Figure 2 As shown, embodiment 2 of the present application provides a work order feature classification system 40 based on natural language processing, including: A first identification module 401 is configured to determine, in response to a work order processing instruction, a salesperson identifier corresponding to a target work order based on target work order data, wherein the work order processing instruction indicates starting to process the target work order data included in the target work order; Acquisition module 402 is configured to acquire a historical position identifier and a current position identifier corresponding to the salesperson identifier; when the historical position identifier and the current position identifier are different, acquire a common ambiguous feature set corresponding to both the current position identifier and the historical position identifier, wherein the common ambiguous feature set includes a plurality of common ambiguous features having the same text characters but different text meanings, wherein the different text meanings indicate that the common ambiguous features have different text meanings when corresponding to different positions; The second recognition module 403 is configured to perform text recognition on the target work order data using a natural language processing algorithm to obtain a plurality of text features; and for each of the text features, if the text feature belongs to the common ambiguous feature set, determine at least one adjacent feature corresponding to the text feature; The processing module 404 is configured to determine an available text meaning corresponding to the text feature according to the text feature and the at least one adjacent feature, so as to classify the text feature in the target work order based on the available text meaning.
[0061] Optionally, the acquisition module 402 is configured to: Determine a current feature identifier set corresponding to the current position identifier and a historical feature identifier set corresponding to the historical position identifier, wherein both the current feature identifier set and the historical feature identifier set include correspondences between a plurality of text characters and text meanings; Determining, based on the current feature identification set and the historical feature identification set, initial ambiguous features of a plurality of text characters having the same characters, and obtaining a plurality of text meanings corresponding to each of the initial ambiguous features; For each of the initial ambiguous features, when the meanings of the multiple texts corresponding to the initial ambiguous feature are different, the initial ambiguous feature is determined to be a common ambiguous feature.
[0062] Optionally, the second recognition module 403 , when determining at least one adjacent feature corresponding to the text feature, is configured to: Determining, based on the plurality of text features, a natural sentence corresponding to the text feature, wherein the natural sentence includes a plurality of text words, and the text word is the text feature or any adjacent entity; The text words are combined in pairs to obtain at least one entity discrimination group, and the adjacent features are obtained based on the entity discrimination group having actual meaning.
[0063] Optionally, the system 40 further includes: Update modules for: Acquire an initial feature identifier set, wherein the initial feature identifier set includes a historical current feature identifier set; Determining whether the plurality of text features are new text features compared to the initial feature identification set; If so, the initial feature identification set is updated according to the new text features to obtain the current feature identification set.
[0064] Optionally, the second recognition module 403 , when determining at least one adjacent feature corresponding to the text feature, is configured to: Determining, based on the plurality of text features, a natural sentence corresponding to the text feature, wherein the natural sentence includes a plurality of text words, and the text word is the text feature or any adjacent entity; Randomly combine the text words in a quantity M to obtain entity discrimination groups corresponding to the quantity M; the quantity M=3,…,N, where N is a positive integer; According to each entity discriminant group corresponding to the quantity M or the quantity 2, the entity discriminant group having actual meaning is used as the adjacent feature.
[0065] Optionally, the natural language processing algorithm runs based on a word vector model algorithm.
[0066] The work order feature classification system based on natural language processing provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effects are similar, and are not described in detail in this embodiment.
[0067] It should be understood that the above-described system embodiments are merely illustrative, and the system of the present application may be implemented in other ways. For example, the division of units / modules in the above-described embodiments is merely a logical functional division, and actual implementations may employ other division methods. For example, multiple units, modules, or components may be combined or integrated into another system, or some features may be omitted or not implemented.
[0068] In addition, unless otherwise specified, the functional units / modules in the various embodiments of the present application may be integrated into a single unit / module, each unit / module may exist physically separately, or two or more units / modules may be integrated together. The aforementioned integrated units / modules may be implemented in the form of hardware or software program modules.
[0069] If the integrated unit / module is implemented in the form of hardware, the hardware may be a digital circuit, an analog circuit, etc. The physical implementation of the hardware structure includes but is not limited to transistors, memristors, etc.
[0070] Figure 3 This is a schematic diagram of the structure of the electronic device provided in this application. Figure 3 As shown, this embodiment 3 provides an electronic device 50 including: at least one processor 501 and a memory 502. Optionally, the electronic device 50 further includes a communication component 503. The processor 501, the memory 502 and the communication component 503 are connected via a bus.
[0071] During the specific implementation process, at least one processor 501 executes the computer execution instructions stored in the memory 502, so that at least one processor 501 executes the work order feature classification method based on natural language processing described in Example 1.
[0072] The specific implementation process of the processor 501 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.
[0073] Unless otherwise specified, the processor 501 may be any appropriate hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC. Unless otherwise specified, the memory 502 may be any appropriate magnetic storage medium or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), and the like.
[0074] If the integrated unit / module is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a memory and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned memory includes: U disk, read-only memory (ROM), random access memory (RAM), mobile hard disk, magnetic disk, or optical disk, etc., various media that can store program code.
[0075] Example 4 of the present application provides a computer-readable storage medium, in which computer execution instructions are stored. When a processor executes the computer execution instructions, the work order feature classification method based on natural language processing described in Example 1 is implemented.
[0076] In the above embodiments, the description of each embodiment has its own emphasis. For parts not described in detail in a particular embodiment, please refer to the relevant description of other embodiments. The technical features of the above embodiments can be combined in any way. To keep the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0077] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0078] It should be understood that the present application is not limited to the exact structure described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A work order feature classification method based on natural language processing, characterized in that: include: In response to a work order processing instruction, determining a salesperson identifier corresponding to the target work order based on target work order data, wherein the work order processing instruction indicates starting to process the target work order data included in the target work order; Obtaining a historical position identifier and a current position identifier corresponding to the salesperson identifier; when the historical position identifier and the current position identifier are different, obtaining a common ambiguous feature set corresponding to both the current position identifier and the historical position identifier, wherein the common ambiguous feature set includes a plurality of common ambiguous features having the same text characters but different text meanings, wherein the different text meanings indicate that the common ambiguous features have different text meanings when corresponding to different positions; Performing text recognition on the target work order data using a natural language processing algorithm to obtain a plurality of text features; and for each of the text features, if the text feature belongs to the common ambiguous feature set, determining at least one adjacent feature corresponding to the text feature; An available text meaning corresponding to the text feature is determined according to the text feature and the at least one adjacent feature, so as to classify the text feature in the target work order based on the available text meaning.
2. The work order feature classification method based on natural language processing according to claim 1 is characterized in that: The step of obtaining the common ambiguous feature set corresponding to the current job identifier and the historical job identifier includes: Determine a current feature identifier set corresponding to the current position identifier and a historical feature identifier set corresponding to the historical position identifier, wherein both the current feature identifier set and the historical feature identifier set include correspondences between a plurality of text characters and text meanings; Determining, based on the current feature identification set and the historical feature identification set, initial ambiguous features of a plurality of text characters having the same characters, and obtaining a plurality of text meanings corresponding to each of the initial ambiguous features; For each of the initial ambiguous features, when the meanings of the multiple texts corresponding to the initial ambiguous feature are different, the initial ambiguous feature is determined to be a common ambiguous feature.
3. The work order feature classification method based on natural language processing according to claim 1 is characterized in that: Determining at least one adjacent feature corresponding to the text feature includes: Determining, based on the plurality of text features, a natural sentence corresponding to the text feature, wherein the natural sentence includes a plurality of text words, and the text word is the text feature or any adjacent entity; The text words are combined in pairs to obtain at least one entity discrimination group, and the adjacent features are obtained based on the entity discrimination group having actual meaning.
4. The work order feature classification method based on natural language processing according to claim 1, characterized in that: Determining at least one adjacent feature corresponding to the text feature includes: Determining, based on the plurality of text features, a natural sentence corresponding to the text feature, wherein the natural sentence includes a plurality of text words, and the text word is the text feature or any adjacent entity; Randomly combine the text words in a quantity M to obtain entity discrimination groups corresponding to the quantity M; the quantity M=3,…,N, where N is a positive integer; According to each entity discriminant group corresponding to the quantity M or the quantity 2, the entity discriminant group having actual meaning is used as the adjacent feature.
5. The work order feature classification method based on natural language processing according to claim 1 is characterized in that: Determining at least one adjacent feature corresponding to the text feature includes: The search interval for adjacent features is defined, with the location of the text feature as the center and expanding to the preset range before and after it. Within the search interval, candidate adjacent features that are semantically related to the text feature are preliminarily screened based on the text's grammatical structure, semantic representation, and logical relationships in the business scenario. Using the semantic analysis model, the semantic relevance between the candidate adjacent features and the text feature is quantitatively evaluated, and the co-occurrence probability between the candidate adjacent features and the text feature is calculated. When the co-occurrence probability is not less than the preset semantic relevance strength threshold, it is determined to be an adjacent feature.
6. The work order feature classification method based on natural language processing according to claim 5 is characterized in that: The candidate adjacent features that are semantically associated with the text feature are initially screened out within the search interval based on the grammatical structure, semantic representation, and logical relationship of the text in the business scenario, including: Perform feature extraction on the text feature and the candidate feature text in the search interval to obtain grammatical features and semantic features of the text, wherein the grammatical features are TF-IDF feature vectors, and the semantic features are semantic vector representations of the text obtained using a pre-trained language model; Calculate the similarity between the text feature and the candidate feature text in terms of grammatical structure, semantic representation, and logical relationship in the business scenario, and obtain grammatical structure similarity, semantic representation similarity, and business logic similarity; Perform weighted fusion on the grammatical structure similarity, semantic representation similarity, business logic similarity and their corresponding weights to obtain the comprehensive similarity score of the candidate feature text; When the comprehensive similarity score is greater than a preset score threshold, the candidate feature text corresponding to the comprehensive similarity score is confirmed as a candidate adjacent feature.
7. The work order feature classification method based on natural language processing according to claim 1 is characterized in that: Before determining the current feature identifier set corresponding to the current position identifier, the method further includes: Acquire an initial feature identifier set, wherein the initial feature identifier set includes a historical current feature identifier set; Determining whether the plurality of text features are new text features compared to the initial feature identification set; If so, the initial feature identification set is updated according to the new text features to obtain the current feature identification set.
8. The work order feature classification method based on natural language processing according to claim 1 is characterized in that: The natural language processing algorithm operates based on a word vector model algorithm.
9. A work order feature classification system based on natural language processing using the work order feature classification method based on natural language processing according to any one of claims 1 to 8, characterized in that: The system comprises: A first identification module is configured to determine, in response to a work order processing instruction, a salesperson identifier corresponding to a target work order based on target work order data, wherein the work order processing instruction indicates starting to process the target work order data included in the target work order; an acquisition module, configured to acquire a historical position identifier and a current position identifier corresponding to the salesperson identifier; and when the historical position identifier and the current position identifier are different, acquire a common ambiguous feature set corresponding to both the current position identifier and the historical position identifier, wherein the common ambiguous feature set includes a plurality of common ambiguous features having the same text characters but different text meanings, wherein the different text meanings indicate that the common ambiguous features have different text meanings when corresponding to different positions; a second recognition module configured to perform text recognition on the target work order data using a natural language processing algorithm to obtain a plurality of text features; and for each of the text features, if the text feature belongs to the common ambiguous feature set, determine at least one adjacent feature corresponding to the text feature; A processing module is configured to determine, based on the text feature and the at least one adjacent feature, an available text meaning corresponding to the text feature, so as to classify the text feature in the target work order based on the available text meaning.
10. An electronic device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the work order feature classification method based on natural language processing as described in any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the work order feature classification method based on natural language processing as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Work order text classification method and device, storage medium and computer equipment
CN114528399A
Work order classification method and device
CN111126842A
Semantic comprehension-based entity recognition method and device, computer equipment and medium
CN112215008A