Engineering resource allocation management recommendation method and system based on big data

By text cleaning and format normalization of project resources project files, word segmentation into word vectors and analyzing project resource information nodes, the problem of inefficient data analysis in project resource project management is solved, efficient resource allocation and project progress control is achieved, and the reliability and user experience of project projects are improved.

CN120470115AInactive Publication Date: 2025-08-12ZHONGYUAN CONSTR MANAGEMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510976174.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-08-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Data analysis in engineering resource project management is inefficient and resource allocation conflicts occur frequently, resulting in low reliability of project progress and quality control, and existing methods are difficult to efficiently connect and interconnect, and changes in resource demand have not been fully considered.

Method used

By obtaining project resources project files, text cleaning and format normalization, word segmentation is used as word units and converted into vectors, word feature vectors are formed based on context features, engineering resource information nodes are analyzed, historical project files are matched, and allocation management strategies are recommended.

Benefits of technology

It improves the analysis accuracy and efficiency of project resource project documents, ensures the accuracy of matching requirements and resources, realizes efficient interconnection of data analysis, and improves the overall reliability and user experience of project projects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470115A_ABST
    Figure CN120470115A_ABST
Patent Text Reader

Abstract

The invention discloses an engineering resource allocation management recommendation method and system based on big data, and the method comprises the steps: obtaining a to-be-analyzed engineering resource project file, carrying out the text cleaning and format normalization of the engineering resource project file, and obtaining a structured project text; performing word segmentation operation on the structured item text to obtain a plurality of word units, and converting the word units into a vector form to obtain word vectors; extracting a context feature of each word vector, and combining the word vector with the corresponding context feature to obtain a word feature vector; according to the character and word feature vectors, a plurality of project resource information nodes are obtained through analysis, and according to the project resource information nodes, corresponding historical project files are matched in a preset project resource database; and retrieving a project resource allocation management strategy corresponding to the historical project file, and outputting and recommending the project resource allocation management strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data processing technology, and in particular to a method and system for recommending engineering resource allocation management based on big data. Background Art

[0002] At present, engineering resource project management involves massive amounts of data, and the data at different stages are complex, making it difficult to process and mine in depth efficiently.

[0003] Currently, existing data analysis methods have many problems. They can only perform preliminary analysis of engineering resource projects, and the analysis of data at each stage is not efficiently interconnected, resulting in low accuracy in matching demand and resources, and difficulty in timely optimization of scheduling management plans. At the same time, because resource scheduling methods mostly rely on manual experience, they do not fully consider changes in resource demand at multiple stages and lack constraint optimization from the data analysis end, which can easily lead to inefficient resource allocation conflicts, thereby affecting the progress and quality control of engineering projects and resulting in low overall reliability. Summary of the Invention

[0004] The present invention provides a method and system for recommending engineering resource allocation management based on big data, so as to solve the technical problems in the prior art of inefficient data analysis of engineering resource projects, easy resource allocation conflicts, and low reliability of engineering resource data management and analysis.

[0005] To solve the above technical problems, an embodiment of the present invention provides a method for engineering resource allocation management recommendation based on big data, comprising: Acquire an engineering resource project file to be analyzed, and perform text cleaning and format normalization on the engineering resource project file to obtain a structured project text; Performing a word segmentation operation on the structured project text to obtain a plurality of word units, and converting the word units into vector form to obtain word vectors; Extracting context features of each word vector and combining the word vector with its corresponding context features to obtain a word feature vector; Analyzing the word feature vector to obtain a plurality of engineering resource information nodes, and matching the corresponding historical engineering files in a preset engineering resource database based on the engineering resource information nodes; wherein the number of the historical engineering files is at least one; The engineering resource allocation management strategy corresponding to the historical engineering file is retrieved, and the engineering resource allocation management strategy is output and recommended.

[0006] As a preferred solution, the process of obtaining the engineering resource project file to be analyzed and performing text cleaning and format normalization on the engineering resource project file to obtain a structured project text specifically includes: In response to a user inputting an engineering resource project file to be analyzed, backing up and storing the engineering resource project file in a preset engineering resource database; Performing text cleaning on the engineering resource project file, and segmenting the text-cleaned engineering resource project file to obtain a plurality of text paragraphs; The format of each text paragraph is normalized to obtain a normalized text paragraph, and the normalized text paragraphs are combined to obtain a structured project text.

[0007] As a preferred solution, the step of normalizing the format of each text paragraph to obtain a normalized text paragraph, and combining the normalized text paragraphs to obtain a structured project text specifically includes: Performing character analysis on each text paragraph and classifying the text paragraph into pure text paragraphs and digital text paragraphs; wherein the text paragraphs of the engineering resource project file include pure text paragraphs and digital text paragraphs; Normalizing the punctuation marks and separators of the plain text paragraph, and uniformly encoding the normalized characters and the plain text paragraph to obtain a normalized plain text paragraph; Performing digital recognition on the digital text paragraph, unifying the format of the recognized digital content, and then uniformly encoding the digital content and the digital text paragraph after the format unification to obtain a normalized digital text paragraph; According to the contextual relationship between the plain text paragraphs and the digital text paragraphs in the engineering resource project file, the normalized plain text paragraphs and the normalized digital text paragraphs are combined to obtain a structured project text.

[0008] As a preferred solution, the word segmentation operation is performed on the structured project text to obtain a number of word units, and the word units are converted into vector form to obtain word vectors, which specifically includes: By using a preset word segmentation model and based on the contextual relationship between the plain text paragraphs and the digital text paragraphs, word segmentation operations are sequentially performed on the plain text paragraphs and the digital text paragraphs in the structured project text, so that during the word segmentation operation, the plain text paragraph or the digital text paragraph currently undergoing word segmentation operation is segmented using the preset word segmentation model, and all word units of the paragraph obtained are input into the preset word segmentation model for self-learning and model updating, so that after the preset word segmentation model is updated, the next plain text paragraph or digital text paragraph is segmented using the updated preset word segmentation model until all paragraphs are segmented to obtain word units corresponding to each paragraph; All word units are converted into vector form through the word vector algorithm to obtain word vectors.

[0009] As a preferred solution, the method for constructing the preset word segmentation model includes: Acquire a historical engineering resource project file; wherein the historical engineering resource project file includes a plurality of annotated word information; Constructing an initial word segmentation model, and inputting each annotated word information in the historical engineering resource project file into the initial word segmentation model for training to obtain a trained word segmentation model; The trained word segmentation model is evaluated for its accuracy, and a preset word segmentation model that passes the evaluation is output.

[0010] As a preferred solution, extracting the context features of each word vector and combining the word vector with its corresponding context features to obtain a word feature vector specifically includes: The word vectors in each paragraph are respectively coded and extracted for contextual features to obtain the upper and lower word vector features corresponding to each word vector, so that for each word vector, its corresponding upper and lower word vector features are used as semantic annotations and combined into a word feature vector, until the word vectors of each paragraph are combined to obtain the corresponding word feature vector; wherein, each word feature vector includes the word vector and its corresponding upper and lower word vector features.

[0011] As a preferred solution, the analysis of the word feature vector to obtain a number of engineering resource information nodes, and matching the corresponding historical engineering files in a preset engineering resource database according to the engineering resource information nodes, specifically includes: Through a preset semantic recognition model, semantic recognition and classification are performed on all word feature vectors to obtain a number of engineering resource information nodes; wherein the engineering resource information nodes include: manpower nodes, equipment nodes, material nodes, funding nodes and time nodes; According to the human node, matching the corresponding first historical engineering file in a preset engineering resource database; According to the device node, matching a corresponding second historical project file in a preset project resource database; Based on the material node, a first similarity evaluation is performed on the material node related information in each historical engineering file in a preset engineering resource database, thereby matching a third historical engineering file whose first similarity is greater than a preset threshold; Based on the funding node, a second similarity evaluation is performed on the funding node related information in each historical project file in the preset project resource database, thereby matching a fourth historical project file whose second similarity is greater than a preset threshold; Establishing a time relationship between the time node and the manpower node, equipment node, material node, and funding node, calculating a time span of the engineering resource project file, and matching a fifth historical engineering file corresponding to the time span in the preset engineering resource database; performing a union process on the first historical project file, the second historical project file, the third historical project file, the fourth historical project file, and the fifth historical project file to obtain a final historical project file; When the number of historical project files obtained by the union is greater than a preset number value, the first historical project file, the second historical project file, the third historical project file, the fourth historical project file and the fifth historical project file are subjected to intersection processing to obtain a final historical project file.

[0012] The present invention also discloses a big data-based engineering resource allocation management recommendation system, comprising: A preprocessing module is used to obtain the engineering resource project file to be analyzed, and perform text cleaning and format normalization on the engineering resource project file to obtain a structured project text; A vector module, configured to perform a word segmentation operation on the structured project text to obtain a plurality of word units, and convert the word units into a vector form to obtain a word vector; a combining module, configured to extract context features of each word vector and combine the word vector with its corresponding context features to obtain a word feature vector; a matching module configured to analyze the word feature vector to obtain a plurality of engineering resource information nodes, and match the corresponding historical engineering files in a preset engineering resource database based on the engineering resource information nodes; wherein the number of the historical engineering files is at least one; The recommendation module is used to retrieve the engineering resource allocation management strategy corresponding to the historical engineering file and output a recommendation for the engineering resource allocation management strategy.

[0013] Accordingly, the present invention also provides a terminal device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements the engineering resource allocation management recommendation method based on big data as described in any one of the above.

[0014] Accordingly, the present invention also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute any one of the big data-based engineering resource allocation management recommendation methods described above.

[0015] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: The technical solution of the present invention obtains the engineering resource project file to be analyzed and directly processes the engineering resource project file, thereby cleaning the file text and normalizing the format, and then performing word segmentation and vector conversion on the structured project text to obtain the corresponding word vector, and combining the context feature to obtain the word feature vector, so that the corresponding word analysis can be performed on the engineering resource project file to ensure that the information of the engineering resource project file can be identified through big data technology to improve the accuracy and efficiency of project file identification. At the same time, various engineering resource information nodes in the engineering resource project file are analyzed through the word feature vector, so that the corresponding historical engineering files can be quickly and accurately matched from the preset engineering resource database, and then the engineering resource allocation management strategy corresponding to the recommended historical engineering file can be retrieved and output, thereby improving the analysis accuracy of the overall project file. At the same time, combined with the historical engineering files in the preset engineering resource database, efficient interconnection and interoperability of data analysis can be achieved to ensure the accuracy of matching demand and resources, thereby greatly improving the efficiency of the entire project file analysis and greatly improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 : A flowchart of a method for recommending engineering resource allocation management based on big data provided by an embodiment of the present invention; Figure 2 : A structural diagram of an engineering resource allocation management recommendation system based on big data provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0018] Example 1

[0019] Please refer to Figure 1 , an embodiment of the present invention provides a method for recommending engineering resource allocation management based on big data, comprising the following steps S101-S105: S101: Acquire an engineering resource project file to be analyzed, and perform text cleaning and format normalization on the engineering resource project file to obtain a structured project text.

[0020] As a preferred solution, the process of obtaining the engineering resource project file to be analyzed and performing text cleaning and format normalization on the engineering resource project file to obtain a structured project text specifically includes: In response to a user inputting an engineering resource project file to be analyzed, backing up and storing the engineering resource project file in a preset engineering resource database; Performing text cleaning on the engineering resource project file, and segmenting the text-cleaned engineering resource project file to obtain a plurality of text paragraphs; The format of each text paragraph is normalized to obtain a normalized text paragraph, and the normalized text paragraphs are combined to obtain a structured project text.

[0021] In this embodiment, after receiving the engineering resource project file to be analyzed input by the user, it needs to be backed up and stored in a preset engineering resource database to ensure that the original file is preserved so that when a failure or problem occurs during subsequent processing, the original data can be restored at any time. It also provides a data basis and guarantee for subsequent analysis operations.

[0022] In this embodiment, the text of the engineering resource project file is cleaned to remove irrelevant information, noise data, and confusingly formatted content. For example, extra spaces, line breaks, and special characters are deleted, and obvious text errors are corrected, thereby making the text content more standardized and unified, which is beneficial for subsequent normalization processing and analysis. Furthermore, the cleaned engineering resource project file is segmented into several relatively independent and semantically complete text paragraphs. This allows the originally large and complex text file to be broken down into smaller, more manageable and analytical units, facilitating subsequent more detailed and in-depth processing of each paragraph and mining the information value contained therein.

[0023] In this embodiment, format normalization is performed on each segmented text paragraph. Since text paragraphs may differ in font, font size, indentation and other formats, by unifying these format elements, each text paragraph is consistent in appearance and data structure, which facilitates subsequent unified data processing and analysis operations, and is also more conducive to the system's recognition and parsing of text. The normalized text paragraphs are combined according to certain logical relationships or established rules to form a structured project text, so that the processed scattered text paragraphs can be integrated into a systematic text content with a clear structure and hierarchy, so that it can better meet the subsequent use requirements in related scenarios such as engineering resource project analysis and application, and facilitate the extraction and utilization of key information, laying the foundation for in-depth engineering resource project analysis and decision support tasks.

[0024] As a preferred solution, the step of normalizing the format of each text paragraph to obtain a normalized text paragraph, and combining the normalized text paragraphs to obtain a structured project text specifically includes: Performing character analysis on each text paragraph and classifying the text paragraph into pure text paragraphs and digital text paragraphs; wherein the text paragraphs of the engineering resource project file include pure text paragraphs and digital text paragraphs; Normalizing the punctuation marks and separators of the plain text paragraph, and uniformly encoding the normalized characters and the plain text paragraph to obtain a normalized plain text paragraph; Performing digital recognition on the digital text paragraph, unifying the format of the recognized digital content, and then uniformly encoding the digital content and the digital text paragraph after the format unification to obtain a normalized digital text paragraph; According to the contextual relationship between the plain text paragraphs and the digital text paragraphs in the engineering resource project file, the normalized plain text paragraphs and the normalized digital text paragraphs are combined to obtain a structured project text.

[0025] In this embodiment, character analysis is performed on each text paragraph, and the text paragraph is classified into pure text paragraphs and digital text paragraphs according to the content characteristics of the text paragraph. Then, the punctuation marks and separators of the pure text paragraph are normalized, and the normalized characters and pure text paragraphs are uniformly encoded. The text content is converted into a standardized representation through encoding, thereby obtaining a normalized pure text paragraph. Digital recognition is performed on the digital text paragraph, and the digital content therein is accurately extracted. The recognized digital content is formatted, for example, the numbers are unified into a specific numerical format, decimal point format, etc., to ensure the consistency of the digital content. The digital content and the digital text paragraph after formatting are uniformly encoded, so that the digital text paragraph can be converted into a standardized representation to obtain a normalized digital text paragraph, so that it can be subsequently integrated with the pure text paragraph.

[0026] In this embodiment, the normalized plain text paragraphs and the normalized digital text paragraphs are combined according to the contextual relationship between the plain text paragraphs and the digital text paragraphs in the engineering resource project file, so that the two types of normalized text paragraphs can be recombined in the original logical order to form a structured project text with a clear structure and hierarchical relationship, so that it can better meet the needs of subsequent engineering resource project analysis and other tasks.

[0027] S102: Perform word segmentation on the structured project text to obtain a number of word units, and convert the word units into vector form to obtain word vectors.

[0028] As a preferred solution, the word segmentation operation is performed on the structured project text to obtain a number of word units, and the word units are converted into vector form to obtain word vectors, which specifically includes: By using a preset word segmentation model and based on the contextual relationship between the plain text paragraphs and the digital text paragraphs, word segmentation operations are sequentially performed on the plain text paragraphs and the digital text paragraphs in the structured project text, so that during the word segmentation operation, the plain text paragraph or the digital text paragraph currently undergoing word segmentation operation is segmented using the preset word segmentation model, and all word units of the paragraph obtained are input into the preset word segmentation model for self-learning and model updating, so that after the preset word segmentation model is updated, the next plain text paragraph or digital text paragraph is segmented using the updated preset word segmentation model until all paragraphs are segmented to obtain word units corresponding to each paragraph; All word units are converted into vector form through the word vector algorithm to obtain word vectors.

[0029] In this embodiment, a preset word segmentation model is used to sequentially segment the plain text paragraphs and the digital text paragraphs in the structured project text according to the contextual relationship between the plain text paragraphs and the digital text paragraphs. During the word segmentation process, the preset word segmentation model is used to segment the plain text paragraph or the digital text paragraph currently undergoing the word segmentation operation. For example, if the current paragraph is a plain text paragraph, the word segmentation model will segment the text in the paragraph into individual word units with independent meanings based on its internal rules and algorithms, such as segmenting "engineering resource project" into word units such as "engineering", "resources", and "project".

[0030] In this embodiment, during the word segmentation operation, all the word units of the paragraph obtained are input into the preset word segmentation model for self-learning and model updating to ensure that the word segmentation model will adjust and optimize its own parameters and rules according to the newly input word units, so that the model can better adapt to new text content and language characteristics, especially for documents such as engineering resource projects, which may involve some new architectural styles, construction methods, different place names, project names, etc., so that the accuracy of subsequent word segmentation can be further improved through model self-learning and updating. For example, when encountering some new word combinations or special expressions during the word segmentation process, the model can update its internal vocabulary and word segmentation rules through self-learning so that similar text content can be processed more accurately in subsequent word segmentation operations. The next pure text paragraph or digital text paragraph is segmented using the updated preset word segmentation model until all paragraphs are segmented. In this process, the preset word segmentation model continuously self-learns and updates according to the new text data, thereby gradually improving the accuracy and adaptability of word segmentation, and ultimately obtaining word units corresponding to each paragraph.

[0031] In this embodiment, a word vector algorithm is used to convert all word units into vector form to obtain word vectors. Semantic information can be represented by converting each word unit token into a vector form. Vector conversion methods include one-hot encoding and word embedding. One-hot encoding represents each token as a high-dimensional one-hot vector whose dimension is the same as the vocabulary size. Word embedding maps tokens into a low-dimensional continuous vector space. Methods such as Word2Vec and GloVe can learn word vectors with semantic and syntactic characteristics based on statistical information such as the co-occurrence of tokens in a corpus.

[0032] As a preferred solution, the method for constructing the preset word segmentation model includes: Acquire a historical engineering resource project file; wherein the historical engineering resource project file includes a plurality of annotated word information; Constructing an initial word segmentation model, and inputting each annotated word information in the historical engineering resource project file into the initial word segmentation model for training to obtain a trained word segmentation model; The trained word segmentation model is evaluated for its accuracy, and a preset word segmentation model that passes the evaluation is output.

[0033] In this embodiment, a historical engineering resource project file is first obtained. The annotated word information contained in the historical engineering resource project file provides annotated data for subsequent model training. Annotated data is a key resource for model training in supervised learning. Next, an initial word segmentation model is constructed. The initial word segmentation model is built based on a specific algorithm and architecture. It can be a statistical word segmentation model (such as an n-gram model), a machine learning-based word segmentation model (such as a conditional random field model), or a deep learning-based word segmentation model (such as a recurrent neural network or Transformer architecture model). The annotated word information in the historical engineering resource project file is input into the initial word segmentation model for training. During the training process, the model learns how to correctly segment the text into word units based on the annotated data. For example, the model will learn that "engineering resources" should be segmented into two word units: "engineering" and "resources," rather than other incorrect segmentation methods.

[0034] In this embodiment, the trained word segmentation model is evaluated for accuracy. The purpose of this evaluation is to determine whether the model's performance meets the expected standards and whether it can accurately segment new text data. This evaluation uses metrics such as accuracy and F1 score. If the model passes the evaluation, it is output as the default word segmentation model.

[0035] S103: Extract the context features of each word vector, and combine the word vector with its corresponding context features to obtain a word feature vector.

[0036] As a preferred solution, extracting the context features of each word vector and combining the word vector with its corresponding context features to obtain a word feature vector specifically includes: The word vectors in each paragraph are respectively coded and extracted for contextual features to obtain the upper and lower word vector features corresponding to each word vector, so that for each word vector, its corresponding upper and lower word vector features are used as semantic annotations and combined into a word feature vector, until the word vectors of each paragraph are combined to obtain the corresponding word feature vector; wherein, each word feature vector includes the word vector and its corresponding upper and lower word vector features.

[0037] In this embodiment, contextual features are encoded and extracted for each word vector in each paragraph. Deep learning models such as recurrent neural networks (RNNs), long short-term memory networks (LSTMs), and gated recurrent units (GRUs) can be used to encode text sequences and obtain the contextual representation of each token. Alternatively, the self-attention mechanism in the Transformer architecture can be used to calculate attention weights between tokens to obtain feature vectors containing global contextual information. This ensures that the semantic information surrounding each word vector is captured. Since the semantics of a word is often influenced by its context, extracting contextual features can more accurately understand the meaning of a word in a specific context.

[0038] Furthermore, the obtained upper and lower word vector features corresponding to each word vector are used as semantic annotations, so that the contextual features can be combined with the vector of the word itself to provide richer semantic information for the word. Then, a word feature vector is formed, and the word vector and its corresponding upper and lower word vector features are combined to form a comprehensive word feature vector. The feature vector contains the semantics of the word itself and the semantic information of its context, which can more comprehensively represent the meaning of the word in the text. The above process is repeated until the word vectors of each paragraph are combined to obtain the corresponding word feature vector. Finally, each word has a feature vector containing its contextual information, which provides a richer semantic basis for subsequent text analysis and processing.

[0039] S104: Analyze and obtain a number of engineering resource information nodes based on the word feature vector, and match corresponding historical engineering files in a preset engineering resource database based on the engineering resource information nodes; wherein the number of the historical engineering files is at least one.

[0040] As a preferred solution, the analysis of the word feature vector to obtain a number of engineering resource information nodes, and matching the corresponding historical engineering files in a preset engineering resource database according to the engineering resource information nodes, specifically includes: Through a preset semantic recognition model, semantic recognition and classification are performed on all word feature vectors to obtain a number of engineering resource information nodes; wherein the engineering resource information nodes include: manpower nodes, equipment nodes, material nodes, funding nodes and time nodes; According to the human node, matching the corresponding first historical engineering file in a preset engineering resource database; According to the device node, matching a corresponding second historical project file in a preset project resource database; Based on the material node, a first similarity evaluation is performed on the material node related information in each historical engineering file in a preset engineering resource database, thereby matching a third historical engineering file whose first similarity is greater than a preset threshold; Based on the funding node, a second similarity evaluation is performed on the funding node related information in each historical project file in the preset project resource database, thereby matching a fourth historical project file whose second similarity is greater than a preset threshold; Establishing a time relationship between the time node and the manpower node, equipment node, material node, and funding node, calculating a time span of the engineering resource project file, and matching a fifth historical engineering file corresponding to the time span in the preset engineering resource database; performing a union process on the first historical project file, the second historical project file, the third historical project file, the fourth historical project file, and the fifth historical project file to obtain a final historical project file; When the number of historical project files obtained by the union is greater than a preset number value, the first historical project file, the second historical project file, the third historical project file, the fourth historical project file and the fifth historical project file are subjected to intersection processing to obtain a final historical project file.

[0041] In this embodiment, the preset semantic recognition model uses a multi-layer perceptron (MLP) to perform nonlinear transformation on the feature vector of the token to learn the semantic discriminant, thereby constructing a semantic recognition model; or a Transformer-based model is used to deeply explore the semantic relationship between tokens and train them through multi-layer stacking of self-attention mechanism and feedforward neural network to achieve accurate understanding and recognition of token semantics.

[0042] In this embodiment, a preset semantic recognition model is used to perform semantic recognition and classification on all word feature vectors, resulting in a number of project resource information nodes, including human resources nodes, equipment nodes, material nodes, funding nodes, and time nodes. For human resources node matching, a first historical project file is matched against the human resources node in a preset project resource database. For equipment nodes, a second historical project file is matched against the preset project resource database. For material node matching, a first similarity evaluation is performed on the material node-related information in each historical project file in the preset project resource database, thereby matching a third historical project file whose first similarity exceeds a preset threshold. For funding node matching, a second similarity evaluation is performed on the funding node-related information in each historical project file in the preset project resource database, thereby matching a fourth historical project file whose second similarity exceeds a preset threshold. A temporal relationship is established between the time node and the human resources node, equipment node, material node, and funding node, and the time span of the project resource project file is calculated. A fifth historical project file corresponding to the time span is then matched against the preset project resource database.

[0043] In this embodiment, a similarity assessment is required for matching material nodes and funding nodes. Therefore, an appropriate similarity assessment algorithm, such as cosine similarity or Jaccard similarity, can be selected. The similarity between the material node or funding node information in each historical project file in the preset project resource database and the current node is calculated. A preset threshold is set. Only when the similarity exceeds this threshold is the match considered successful. For example, setting the cosine similarity threshold to 0.8 means that only when the similarity exceeds 0.8 will the corresponding historical project file be considered a match result.

[0044] In this embodiment, the relationships between time nodes and labor nodes, equipment nodes, material nodes, and funding nodes are analyzed. Based on the established time relationships, the time span of the engineering resource project file is then calculated. For example, the time span can be obtained by calculating the difference between the start time and the end time. Finally, a fifth historical engineering file that matches the calculated time span is searched within a preset engineering resource database. This can be achieved by comparing the time spans of the historical engineering files in the database with the target time span. Alternatively, a certain error range can be set to increase matching flexibility.

[0045] In this embodiment, a set operation or database query operation is used to merge the first, second, third, fourth, and fifth historical project files together to obtain a preliminary set of historical project files. For example, a UNION operation can be used in the database to implement the union process. When the number of historical project files obtained by the union exceeds a preset value, a set operation or database query operation is used to perform an intersection process on the historical project files. For example, an INTERSECT operation can be used in the database to implement the intersection process, thereby ensuring that historical project files that simultaneously meet multiple conditions are screened out, thereby improving the quality and accuracy of the matching results.

[0046] S105: Retrieve the engineering resource allocation management strategy corresponding to the historical engineering file, and output and recommend the engineering resource allocation management strategy.

[0047] In this embodiment, the corresponding engineering resource allocation management strategy is retrieved based on the matched historical engineering files. Engineering resource allocation management strategies are based on experience gained from past projects and include resource allocation methods for manpower, equipment, materials, funds, and time. The retrieved engineering resource allocation management strategies are output and recommended to ensure that effective resource allocation management experience is provided to current engineering resource projects for reference and application, thereby improving the rationality and efficiency of resource allocation.

[0048] Specifically, to search for project resource allocation management strategies, one can first index historical project files in a pre-set project resource database, specifically targeting the project resource allocation management strategy. Based on the identifiers or keywords of the matched historical project files, the corresponding project resource allocation management strategies can be found in the index. For example, if the matched historical project file is about a construction project, resource allocation management strategies related to construction projects are retrieved. The recommended project resource allocation management strategies are then optimized to ensure practicality and operability, and presented to users through appropriate output formats (such as reports or electronic documents). Furthermore, supplementary explanations or case studies can be provided to help users better understand and apply these strategies.

[0049] The implementation of the above embodiment has the following effects: The technical solution of the present invention obtains the engineering resource project file to be analyzed and directly processes the engineering resource project file, thereby cleaning the file text and normalizing the format, and then performing word segmentation and vector conversion on the structured project text to obtain the corresponding word vector, and combining the context feature to obtain the word feature vector, so that the corresponding word analysis can be performed on the engineering resource project file to ensure that the information of the engineering resource project file can be identified through big data technology to improve the accuracy and efficiency of project file identification. At the same time, various engineering resource information nodes in the engineering resource project file are analyzed through the word feature vector, so that the corresponding historical engineering files can be quickly and accurately matched from the preset engineering resource database, and then the engineering resource allocation management strategy corresponding to the recommended historical engineering file can be retrieved and output, thereby improving the analysis accuracy of the overall project file. At the same time, combined with the historical engineering files in the preset engineering resource database, efficient interconnection and interoperability of data analysis can be achieved to ensure the accuracy of matching demand and resources, thereby greatly improving the efficiency of the entire project file analysis and greatly improving the user experience.

[0050] Example 2

[0051] See also Figure 2 , which is a big data-based engineering resource allocation management recommendation system provided by the present invention, including: The pre-processing module 201 is used to obtain the engineering resource project file to be analyzed, and perform text cleaning and format normalization on the engineering resource project file to obtain a structured project text; A vector module 202 is configured to perform a word segmentation operation on the structured project text to obtain a plurality of word units, and convert the word units into vector form to obtain word vectors; A combining module 203 is configured to extract context features of each word vector and combine the word vector with its corresponding context features to obtain a word feature vector; A matching module 204 is configured to analyze the word feature vector to obtain a plurality of engineering resource information nodes, and match the corresponding historical engineering files in a preset engineering resource database based on the engineering resource information nodes; wherein the number of the historical engineering files is at least one; The recommendation module 205 is configured to retrieve the engineering resource allocation management strategy corresponding to the historical engineering file and output a recommendation for the engineering resource allocation management strategy.

[0052] As a preferred solution, the process of obtaining the engineering resource project file to be analyzed and performing text cleaning and format normalization on the engineering resource project file to obtain a structured project text specifically includes: In response to a user inputting an engineering resource project file to be analyzed, backing up and storing the engineering resource project file in a preset engineering resource database; Performing text cleaning on the engineering resource project file, and segmenting the text-cleaned engineering resource project file to obtain a plurality of text paragraphs; The format of each text paragraph is normalized to obtain a normalized text paragraph, and the normalized text paragraphs are combined to obtain a structured project text.

[0053] As a preferred solution, the step of normalizing the format of each text paragraph to obtain a normalized text paragraph, and combining the normalized text paragraphs to obtain a structured project text specifically includes: Performing character analysis on each text paragraph and classifying the text paragraph into pure text paragraphs and digital text paragraphs; wherein the text paragraphs of the engineering resource project file include pure text paragraphs and digital text paragraphs; Normalizing the punctuation marks and separators of the plain text paragraph, and uniformly encoding the normalized characters and the plain text paragraph to obtain a normalized plain text paragraph; Performing digital recognition on the digital text paragraph, unifying the format of the recognized digital content, and then uniformly encoding the digital content and the digital text paragraph after the format unification to obtain a normalized digital text paragraph; According to the contextual relationship between the plain text paragraphs and the digital text paragraphs in the engineering resource project file, the normalized plain text paragraphs and the normalized digital text paragraphs are combined to obtain a structured project text.

[0054] As a preferred solution, the word segmentation operation is performed on the structured project text to obtain a number of word units, and the word units are converted into vector form to obtain word vectors, which specifically includes: By using a preset word segmentation model and based on the contextual relationship between the plain text paragraphs and the digital text paragraphs, word segmentation operations are sequentially performed on the plain text paragraphs and the digital text paragraphs in the structured project text, so that during the word segmentation operation, the plain text paragraph or the digital text paragraph currently undergoing word segmentation operation is segmented using the preset word segmentation model, and all word units of the paragraph obtained are input into the preset word segmentation model for self-learning and model updating, so that after the preset word segmentation model is updated, the next plain text paragraph or digital text paragraph is segmented using the updated preset word segmentation model until all paragraphs are segmented to obtain word units corresponding to each paragraph; All word units are converted into vector form through the word vector algorithm to obtain word vectors.

[0055] As a preferred solution, the method for constructing the preset word segmentation model includes: Acquire a historical engineering resource project file; wherein the historical engineering resource project file includes a plurality of annotated word information; Constructing an initial word segmentation model, and inputting each annotated word information in the historical engineering resource project file into the initial word segmentation model for training to obtain a trained word segmentation model; The trained word segmentation model is evaluated for its accuracy, and a preset word segmentation model that passes the evaluation is output.

[0056] As a preferred solution, extracting the context features of each word vector and combining the word vector with its corresponding context features to obtain a word feature vector specifically includes: The word vectors in each paragraph are respectively coded and extracted for contextual features to obtain the upper and lower word vector features corresponding to each word vector, so that for each word vector, its corresponding upper and lower word vector features are used as semantic annotations and combined into a word feature vector, until the word vectors of each paragraph are combined to obtain the corresponding word feature vector; wherein, each word feature vector includes the word vector and its corresponding upper and lower word vector features.

[0057] As a preferred solution, the analysis of the word feature vector to obtain a number of engineering resource information nodes, and matching the corresponding historical engineering files in a preset engineering resource database according to the engineering resource information nodes, specifically includes: Through a preset semantic recognition model, semantic recognition and classification are performed on all word feature vectors to obtain a number of engineering resource information nodes; wherein the engineering resource information nodes include: manpower nodes, equipment nodes, material nodes, funding nodes and time nodes; According to the human node, matching the corresponding first historical engineering file in a preset engineering resource database; According to the device node, matching a corresponding second historical project file in a preset project resource database; Based on the material node, a first similarity evaluation is performed on the material node related information in each historical engineering file in a preset engineering resource database, thereby matching a third historical engineering file whose first similarity is greater than a preset threshold; Based on the funding node, a second similarity evaluation is performed on the funding node related information in each historical project file in the preset project resource database, thereby matching a fourth historical project file whose second similarity is greater than a preset threshold; Establishing a time relationship between the time node and the manpower node, equipment node, material node, and funding node, calculating a time span of the engineering resource project file, and matching a fifth historical engineering file corresponding to the time span in the preset engineering resource database; performing a union process on the first historical project file, the second historical project file, the third historical project file, the fourth historical project file, and the fifth historical project file to obtain a final historical project file; When the number of historical project files obtained by the union is greater than a preset number value, the first historical project file, the second historical project file, the third historical project file, the fourth historical project file and the fifth historical project file are subjected to intersection processing to obtain a final historical project file.

[0058] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0059] The implementation of the above embodiment has the following effects: The technical solution of the present invention obtains the engineering resource project file to be analyzed and directly processes the engineering resource project file, thereby cleaning the file text and normalizing the format, and then performing word segmentation and vector conversion on the structured project text to obtain the corresponding word vector, and combining the context feature to obtain the word feature vector, so that the corresponding word analysis can be performed on the engineering resource project file to ensure that the information of the engineering resource project file can be identified through big data technology to improve the accuracy and efficiency of project file identification. At the same time, various engineering resource information nodes in the engineering resource project file are analyzed through the word feature vector, so that the corresponding historical engineering files can be quickly and accurately matched from the preset engineering resource database, and then the engineering resource allocation management strategy corresponding to the recommended historical engineering file can be retrieved and output, thereby improving the analysis accuracy of the overall project file. At the same time, combined with the historical engineering files in the preset engineering resource database, efficient interconnection and interoperability of data analysis can be achieved to ensure the accuracy of matching demand and resources, thereby greatly improving the efficiency of the entire project file analysis and greatly improving the user experience.

[0060] Example 3

[0061] Accordingly, the present invention also provides a terminal device, comprising: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements the engineering resource allocation management recommendation method based on big data as described in any one of the above embodiments.

[0062] The terminal device of this embodiment includes: a processor, a memory, and a computer program and computer instructions stored in the memory and capable of running on the processor. When the processor executes the computer program, each step in the above embodiment 1 is implemented, such as Figure 1 Alternatively, when the processor executes the computer program, the functions of each module / unit in the above-mentioned device embodiment, such as the matching module 204, are implemented.

[0063] Exemplarily, the computer program can be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments that can perform specific functions, and the instruction segments are used to describe the execution process of the computer program in the terminal device. For example, the matching module 204 is used to analyze the word feature vector to obtain a number of engineering resource information nodes, and match the corresponding historical engineering files in the preset engineering resource database based on the engineering resource information nodes.

[0064] The terminal device may be a computing device such as a desktop computer, laptop, PDA, or cloud server. The terminal device may include, but is not limited to, a processor and memory. Those skilled in the art will appreciate that the schematic diagram is merely an example of a terminal device and does not limit the terminal device. The terminal device may include more or fewer components than shown, or a combination of certain components, or different components. For example, the terminal device may also include input / output devices, network access devices, buses, and the like.

[0065] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor. The processor is the control center of the terminal device and connects various parts of the entire terminal device using various interfaces and lines.

[0066] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the terminal device by running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store an operating system, at least one application required for a function, etc.; the data storage area may store data created based on the use of the mobile terminal, etc. In addition, the memory may include high-speed random access memory and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0067] If the module / unit integrated into the terminal device is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention can implement all or part of the process steps in the above-mentioned method embodiments by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content of the computer-readable medium can be appropriately increased or decreased based on the requirements of legislation and patent practice in a jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0068] Example 4

[0069] Accordingly, the present invention also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the engineering resource allocation management recommendation method based on big data as described in any one of the above embodiments.

[0070] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A method for recommending engineering resource allocation management based on big data, characterized in that: include: Acquire an engineering resource project file to be analyzed, and perform text cleaning and format normalization on the engineering resource project file to obtain a structured project text; Performing a word segmentation operation on the structured project text to obtain a plurality of word units, and converting the word units into vector form to obtain word vectors; Extracting context features of each word vector and combining the word vector with its corresponding context features to obtain a word feature vector; Analyzing the word feature vector to obtain a plurality of engineering resource information nodes, and matching the corresponding historical engineering files in a preset engineering resource database based on the engineering resource information nodes; wherein the number of the historical engineering files is at least one; The engineering resource allocation management strategy corresponding to the historical engineering file is retrieved, and the engineering resource allocation management strategy is output and recommended.

2. The method for recommending engineering resource allocation management based on big data according to claim 1, characterized in that: The step of obtaining the engineering resource project file to be analyzed and performing text cleaning and format normalization on the engineering resource project file to obtain a structured project text specifically includes: In response to a user inputting an engineering resource project file to be analyzed, backing up and storing the engineering resource project file in a preset engineering resource database; Performing text cleaning on the engineering resource project file, and segmenting the text-cleaned engineering resource project file to obtain a plurality of text paragraphs; The format of each text paragraph is normalized to obtain a normalized text paragraph, and the normalized text paragraphs are combined to obtain a structured project text.

3. The method for recommending engineering resource allocation management based on big data according to claim 2, characterized in that: Normalizing the format of each text paragraph to obtain a normalized text paragraph, and combining the normalized text paragraphs to obtain a structured project text, specifically includes: Performing character analysis on each text paragraph and classifying the text paragraph into pure text paragraphs and digital text paragraphs; wherein the text paragraphs of the engineering resource project file include pure text paragraphs and digital text paragraphs; Normalizing the punctuation marks and separators of the plain text paragraph, and uniformly encoding the normalized characters and the plain text paragraph to obtain a normalized plain text paragraph; Performing digital recognition on the digital text paragraph, unifying the format of the recognized digital content, and then uniformly encoding the digital content and the digital text paragraph after the format unification to obtain a normalized digital text paragraph; According to the contextual relationship between the plain text paragraphs and the digital text paragraphs in the engineering resource project file, the normalized plain text paragraphs and the normalized digital text paragraphs are combined to obtain a structured project text.

4. The method for recommending engineering resource allocation management based on big data according to claim 3, characterized in that: The word segmentation operation is performed on the structured project text to obtain a plurality of word units, and the word units are converted into vector form to obtain word vectors, specifically including: By using a preset word segmentation model and based on the contextual relationship between the plain text paragraphs and the digital text paragraphs, word segmentation operations are sequentially performed on the plain text paragraphs and the digital text paragraphs in the structured project text, so that during the word segmentation operation, the plain text paragraph or the digital text paragraph currently undergoing word segmentation operation is segmented using the preset word segmentation model, and all word units of the paragraph obtained are input into the preset word segmentation model for self-learning and model updating, so that after the preset word segmentation model is updated, the next plain text paragraph or digital text paragraph is segmented using the updated preset word segmentation model until all paragraphs are segmented to obtain word units corresponding to each paragraph; All word units are converted into vector form through the word vector algorithm to obtain word vectors.

5. The method for recommending engineering resource allocation management based on big data according to claim 4, characterized in that: The method for constructing the preset word segmentation model includes: Acquire a historical engineering resource project file; wherein the historical engineering resource project file includes a plurality of annotated word information; Constructing an initial word segmentation model, and inputting each annotated word information in the historical engineering resource project file into the initial word segmentation model for training to obtain a trained word segmentation model; The trained word segmentation model is evaluated for its accuracy, and a preset word segmentation model that passes the evaluation is output.

6. The method for recommending engineering resource allocation management based on big data according to claim 4, characterized in that: The process of extracting context features of each word vector and combining the word vector with its corresponding context features to obtain a word feature vector specifically includes: The word vectors in each paragraph are respectively coded and extracted for contextual features to obtain the upper and lower word vector features corresponding to each word vector, so that for each word vector, its corresponding upper and lower word vector features are used as semantic annotations and combined into a word feature vector, until the word vectors of each paragraph are combined to obtain the corresponding word feature vector; wherein, each word feature vector includes the word vector and its corresponding upper and lower word vector features.

7. The method for recommending engineering resource allocation management based on big data according to any one of claims 1 to 6, characterized in that: The step of analyzing the word feature vector to obtain a plurality of engineering resource information nodes and matching corresponding historical engineering files in a preset engineering resource database according to the engineering resource information nodes specifically includes: Through a preset semantic recognition model, semantic recognition and classification are performed on all word feature vectors to obtain a number of engineering resource information nodes; wherein the engineering resource information nodes include: manpower nodes, equipment nodes, material nodes, funding nodes and time nodes; According to the human node, matching the corresponding first historical engineering file in a preset engineering resource database; According to the device node, matching a corresponding second historical project file in a preset project resource database; Based on the material node, a first similarity evaluation is performed on the material node related information in each historical engineering file in a preset engineering resource database, thereby matching a third historical engineering file whose first similarity is greater than a preset threshold; Based on the funding node, a second similarity evaluation is performed on the funding node related information in each historical project file in the preset project resource database, thereby matching a fourth historical project file whose second similarity is greater than a preset threshold; Establishing a time relationship between the time node and the manpower node, equipment node, material node, and funding node, calculating a time span of the engineering resource project file, and matching a fifth historical engineering file corresponding to the time span in the preset engineering resource database; performing a union process on the first historical project file, the second historical project file, the third historical project file, the fourth historical project file, and the fifth historical project file to obtain a final historical project file; When the number of historical project files obtained by the union is greater than a preset number value, the first historical project file, the second historical project file, the third historical project file, the fourth historical project file and the fifth historical project file are subjected to intersection processing to obtain a final historical project file.

8. A big data-based engineering resource allocation management recommendation system, characterized in that: include: A preprocessing module is used to obtain the engineering resource project file to be analyzed, and perform text cleaning and format normalization on the engineering resource project file to obtain a structured project text; A vector module, configured to perform a word segmentation operation on the structured project text to obtain a plurality of word units, and convert the word units into a vector form to obtain a word vector; a combining module, configured to extract context features of each word vector and combine the word vector with its corresponding context features to obtain a word feature vector; a matching module configured to analyze the word feature vector to obtain a plurality of engineering resource information nodes, and match the corresponding historical engineering files in a preset engineering resource database based on the engineering resource information nodes; wherein the number of the historical engineering files is at least one; The recommendation module is used to retrieve the engineering resource allocation management strategy corresponding to the historical engineering file and output a recommendation for the engineering resource allocation management strategy.

9. A terminal device, characterized in that: The invention comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the method for recommending engineering resource allocation management based on big data as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the engineering resource allocation management recommendation method based on big data as described in any one of claims 1 to 7.