Enterprise user portrait understanding method, system and equipment and storage medium
By establishing a prompt word template library and external information to filter supplementary features, the problem of inaccurate corporate portraits caused by the lack of targeted pre-trained models is solved, and higher-quality corporate feature extraction and portrait generation are achieved.
Patent Information
- Application Number
- CN202510328203.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, the pre-trained model lacks targeting, resulting in the generated corporate portraits lacking accuracy, and it is difficult to extract high-quality corporate characteristics from diversified and complex corporate business text data.
By establishing a library of prompt word templates, selecting prompt word templates corresponding to the business text data type for feature extraction, modifying the preset model in combination with the initial enterprise portrait, and using the relevance of external business information to filter and supplement feature generation, enriching the feature dimensions of the enterprise portrait.
It improves the accuracy and pertinence of corporate portraits, ensures the accuracy of feature extraction and the adaptability of models, and the generated corporate portraits are more comprehensive and accurate.
Smart Images

Figure CN120256633A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of enterprise portrait technology, and specifically to an enterprise user portrait understanding method, system, device and storage medium. Background Art
[0002] Enterprise user portraits are an important part of enterprise informatization construction. By accurately describing enterprise characteristics, they can provide data support for enterprise management, business decision-making and development planning. However, enterprise business text data is usually diverse and complex. How to accurately extract enterprise characteristics from these unstructured text data and generate high-quality enterprise portraits has become a technical problem that needs to be solved in this field.
[0003] Currently, fixed pre-trained models are usually used for feature extraction, that is, a pre-trained general model is used to directly process and analyze the business text data of an enterprise to generate a corporate portrait. Although the above method can achieve basic corporate feature extraction, the lack of targeted optimization of specific corporate features in the pre-trained model leads to the lack of accuracy of the generated corporate portrait. Summary of the invention
[0004] The present application provides a method, system, device and storage medium for understanding enterprise user portraits, which are used to improve the accuracy of generated enterprise portraits.
[0005] In a first aspect, the present application provides a method for understanding enterprise user portraits, the method comprising: obtaining enterprise business text data, selecting a prompt word template corresponding to the business text data type from a prompt word template library; performing feature extraction on the business text data through the prompt word template to obtain initial feature data; inputting the initial feature data into a preset model to generate an initial enterprise portrait of the enterprise, and correcting the preset model through the initial enterprise portrait to obtain a target model; retrieving multiple external business information of the enterprise, calculating the correlation value of each external business information with the initial enterprise portrait, and inputting the external business information whose correlation value is greater than a preset threshold into the target model to obtain supplementary features; inputting the initial enterprise portrait and the supplementary features into the target model to generate an enterprise portrait label for the enterprise.
[0006] By adopting the above technical solution, by selecting the prompt word template corresponding to the business text data type in the prompt word template library for feature extraction, accurate processing of different types of business data can be achieved; the preset model is corrected through the initial enterprise portrait to obtain the target model, thereby improving the pertinence of the model; and combined with the relevance screening of external business information and the generation of supplementary features, the feature dimensions of the enterprise portrait are enriched, and the accuracy of the generated enterprise portrait is improved.
[0007] Optionally, before selecting the prompt template corresponding to the business text data type from the prompt template library, it further includes: obtaining the business text data set of the enterprise, classifying the business text data set to obtain multiple data sets; selecting keywords corresponding to the text features of each data set from a preset keyword library; annotating each data set according to the keywords to obtain various types of business text data; generating corresponding prompt templates according to each type of business text data, storing the prompt templates in the prompt template library, and establishing a mapping relationship between each type of business text data and the corresponding prompt template.
[0008] By adopting the above technical solution, through classifying the business text data set and using the keywords in the preset keyword library to annotate different data sets, the refined classification of business text data is realized; further, corresponding prompt templates are generated based on each type of business text data after annotation, and a mapping relationship between the data type and the template is established, so that the prompt template library can accurately match different types of business text data, improving the pertinence and accuracy of subsequent feature extraction.
[0009] Optionally, the feature extraction of the business text data through the prompt template to obtain initial feature data includes: obtaining the feature extraction rules in the prompt template, where the feature extraction rules include text classification rules and entity recognition rules; performing classification processing on the business text data according to the text classification rules to obtain classification features; performing entity recognition on the business text data according to the entity recognition rules to obtain entity features; performing vectorization processing on the classification features to obtain classification feature vectors; performing vectorization processing on the entity features to obtain entity feature vectors; splicing the classification feature vectors and the entity feature vectors to obtain initial feature data.
[0010] By adopting the above technical solution, the business text data is processed respectively through the text classification rules and entity recognition rules in the feature extraction rules to obtain classification features and entity features, and these two types of features are respectively vectorized and then spliced to form initial feature data containing multi-dimensional information, realizing the comprehensive capture and structured expression of enterprise features, and providing a richer and more accurate feature basis for the subsequent generation of enterprise portraits.
[0011] Optionally, before inputting the initial feature data into a preset model to generate an initial enterprise portrait of the enterprise, it further includes: obtaining an annotated sample data set, dividing the annotated sample data set into a training sample set and a validation sample set according to a preset ratio; constructing an initial preset model based on the training sample set; inputting the validation sample set into the initial preset model for validation to obtain a validation result; optimizing the initial preset model according to the validation result until the validation result meets a preset condition to obtain a preset model.
[0012] By adopting the above technical solution, by dividing the annotated sample data set into a training sample set and a validation sample set, constructing an initial preset model using the training sample set, and iteratively optimizing the model based on the validation result of the validation sample set until the preset condition is met, it ensures that the preset model has good generalization ability and stability, laying a reliable model foundation for subsequent generation of the initial enterprise portrait.
[0013] Optionally, the method of correcting the preset model with the initial enterprise portrait to obtain a target model includes: obtaining the feature parameters in the initial enterprise portrait and using the feature parameters as training data; iteratively adjusting the model parameters of the preset model based on the training data to obtain multiple groups of candidate model parameters; calculating the performance evaluation values of each candidate model parameter, and using the candidate model parameter with the largest performance evaluation value as the target parameter; updating the model parameters of the preset model to the target parameter to obtain a target model.
[0014] By adopting the above technical solution, by using the feature parameters in the initial enterprise portrait as training data, iteratively adjusting the preset model to obtain multiple groups of candidate model parameters, and selecting the optimal target parameter to update the model through the performance evaluation value, it realizes the adaptive optimization of the model, enables the target model to better adapt to the characteristics of a specific enterprise, and improves the pertinence and accuracy of the model.
[0015] Optionally, the method of inputting the external business information with a relevance value greater than a preset threshold into the target model to obtain supplementary features includes: performing text preprocessing on the external business information with a relevance value greater than the preset threshold to obtain preprocessed text; extracting a feature extraction layer from the target model and inputting the preprocessed text into the feature extraction layer to obtain a feature vector; performing feature fusion on the feature vector based on a preset fusion rule to obtain a fusion feature; inputting the fusion feature into the feature mapping layer of the target model to obtain supplementary features.
[0016] By adopting the above technical solution, through text preprocessing of the screened external business information, obtaining feature vectors using the feature extraction layer of the target model, then performing feature fusion based on a preset fusion rule, and finally generating supplementary features through the feature mapping layer, the effective utilization of external business information and the in-depth integration of features are achieved, enabling the finally generated enterprise portrait to contain richer external information dimensions.
[0017] Optionally, the step of inputting the initial enterprise portrait and the supplementary features into the target model to generate the enterprise portrait label of the enterprise includes: merging the initial enterprise portrait and the supplementary features to obtain merged features; performing normalization processing on the merged features to obtain normalized features; inputting the normalized features into the classification layer of the target model to obtain prediction probability values; and mapping the prediction probability values to the enterprise portrait label corresponding to the enterprise according to a preset label mapping rule.
[0018] By adopting the above technical solution, through merging and normalizing the initial enterprise portrait and the supplementary features to obtain normalized features, calculating prediction probability values using the classification layer of the target model, and finally generating enterprise portrait labels based on the label mapping rule, the effective integration of internal and external features is achieved, enabling the finally generated enterprise portrait labels to not only maintain data consistency but also reflect the comprehensive expression of multi-dimensional features.
[0019] In a second aspect, the present application provides an enterprise user portrait understanding system, which includes: an acquisition module, an extraction module, a first input module, a calculation module, and a second input module; where, The acquisition module is used to acquire enterprise business text data and select a prompt word template corresponding to the type of the business text data from a prompt word template library; the extraction module is used to extract features from the business text data through the prompt word template to obtain initial feature data; the first input module is used to input the initial feature data into a preset model to generate the initial enterprise portrait of the enterprise, and correct the preset model through the initial enterprise portrait to obtain a target model; the calculation module is used to retrieve multiple external business information of the enterprise, calculate the relevance value between each external business information and the initial enterprise portrait, and input the external business information with the relevance value greater than a preset threshold into the target model to obtain supplementary features; the second input module is used to input the initial enterprise portrait and the supplementary features into the target model to generate the enterprise portrait label of the enterprise.
[0020] In a third aspect, the present application provides an electronic device, which adopts the following technical solution: It includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. The user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory, so that the electronic device executes a computer program of any one of the above enterprise user portrait understanding methods.
[0021] In a fourth aspect, the present application provides a computer-readable storage medium, which adopts the following technical solution: It stores a computer program that can be loaded and executed by a processor for any one of the above enterprise user portrait understanding methods.
[0022] In summary, the present application includes at least one of the following beneficial technical effects: By selecting a prompt word template corresponding to the business text data type from the prompt word template library for feature extraction, precise processing of different types of business data can be achieved; by correcting a preset model with an initial enterprise portrait to obtain a target model, the pertinence of the model is improved; and by combining the relevance screening of external business information and the generation of supplementary features, the feature dimension of the enterprise portrait is enriched, and the accuracy of the generated enterprise portrait is improved. Description of the Drawings
[0023] Figure 1 is a flowchart of an enterprise user portrait understanding method provided by an embodiment of the present application; Figure 2 is a structural schematic diagram of an enterprise user portrait understanding system provided by an embodiment of the present application; Figure 3 is a structural schematic diagram of an electronic device provided by an embodiment of the present application.
[0024] Description of the reference numerals: 1000, electronic device; 1001, processor; 1002, communication bus; 1003, user interface; 1004, network interface; 1005, memory. Detailed Embodiments
[0025] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.
[0026] In the description of the embodiments of the present application, words such as "exemplary", "for example", or "for illustration" are used to denote examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary", "for example", or "for illustration" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example", or "for illustration" is intended to present relevant concepts in a specific manner.
[0027] Figure 1 is a schematic flowchart of a method for understanding enterprise user portraits provided by the embodiments of the present application. As Figure 1 shown, the method includes S101 - S105: S101, Obtain enterprise business text data, and select a prompt word template corresponding to the type of the business text data from the prompt word template library.
[0028] In this embodiment, first, enterprise business text data is obtained. The enterprise business text data includes, but is not limited to, the enterprise's patent documents, bidding information, market activity data, industrial and commercial registration information, legal information, financial and tax information, etc. Since different types of business text data have different text characteristics and data structures, for example, patent documents usually contain fixed chapters such as technical fields, technical backgrounds, and invention contents, while bidding information contains different content structures such as project descriptions, bidding requirements, and scoring criteria, it is necessary to select corresponding processing strategies for different types of business text data.
[0029] To improve the accuracy and efficiency of text processing, a prompt word template library is pre - established in this embodiment. The prompt word template library stores prompt word templates corresponding to different types of business text data. Each prompt word template contains processing rules and extraction strategies for specific types of business text. For example, for business text data of the patent document type, the corresponding prompt word template contains extraction rules for identifying technical features, inventors, patent classifications, etc.; for business text data of the bidding information type, the corresponding prompt word template contains processing rules such as project category judgment, amount extraction, and technical requirement identification.
[0030] After obtaining the enterprise business text data, the system will automatically identify the type of the business text data. Specifically, the text type can be judged by analyzing information such as the format features, keyword distribution, and chapter structure of the text. After identifying the text type, the system will select a prompt word template corresponding to this type from the prompt word template library. For example, when a text is identified as a patent document, the system will select the prompt word template corresponding to the patent document, which contains specific prompt words and processing rules such as "Please identify the technical field, the technical problems to be solved, and the specific implementation steps of the technical solution" in the text.
[0031] By using pre-customized prompt templates, key information in different types of business texts can be captured more accurately, avoiding information loss or incorrect extraction problems that may be caused by general processing methods. At the same time, the use of prompt templates also makes the processing process more standardized and normalized, which is conducive to maintaining the consistency of extraction results. For example, when processing the patent documents of a technology company, by using a special prompt template for patent documents, key information such as the core algorithm features and technological innovation points in the field of image recognition technology of the company were successfully extracted, providing accurate basic data for subsequent enterprise portrait construction.
[0032] On the basis of the above embodiments, as an optional implementation manner, before selecting the prompt template corresponding to the business text data type from the prompt template library in S101, it specifically further includes S11-S14: S11, obtain the business text data set of the enterprise, classify the business text data set, and obtain multiple data sets.
[0033] The business text data set is a collection of various text materials generated by the enterprise in the daily operation process, and may contain text data in different formats and from different sources. Only by reasonably classifying and annotating these data can targeted prompt templates be generated, thereby improving the accuracy of subsequent feature extraction.
[0034] After obtaining the business text data set, the system classifies the data set by using a method combining rules and machine learning. The classification process first makes a preliminary classification based on the format features (such as document structure, chapter arrangement) and content features (such as professional terms, industry terms) of the text, and then uses a pre-trained text classification model to optimize and adjust the classification results. For example, the system may classify the text data into multiple data sets such as patent literature category, R & D report category, market analysis category, financial statement category, etc. This classification method ensures that similar processing strategies can be obtained for the same type of text data.
[0035] S12, select keywords corresponding to the text features of each data set from the preset keyword library.
[0036] The preset keyword library is a dictionary containing professional terms and key expressions in various business fields, which is classified and organized according to different business fields and text types. The system selects corresponding keywords from the preset keyword library according to the text features of each data set. For example, for the patent literature data set, the system will select keywords related to the technical field, invention purpose, technical solution, etc.; for the market analysis data set, keywords related to market size, competitive situation, development trend, etc. will be selected. These keywords will serve as the basis for subsequent text annotation.
[0037] S13. Annotate each dataset according to the keywords to obtain various types of business text data.
[0038] S14. Generate corresponding prompt templates according to various types of business text data, store the prompt templates in the prompt template library, and establish a mapping relationship between various types of business text data and the corresponding prompt templates.
[0039] The text annotation process is carried out in a semi - automated manner. The system first automatically annotates the text in each dataset using the selected keywords. The annotation process includes keyword matching, semantic analysis, and context understanding. For example, when processing a technical patent document, the system will annotate information such as "Technical Field: Image Recognition", "Application Scenario: Intelligent Security", etc. To improve the annotation quality, the system also uses manual review to verify and supplement the automatic annotation results to ensure the accuracy and integrity of the annotation.
[0040] Based on the annotated business text data, the system starts to generate prompt templates. The generation process of the prompt templates takes into account the characteristics of the text type and the distribution law of the annotation information. For example, for R & D - type texts, the generated prompt templates may contain prompt messages such as "Please identify the core technical features in the text", "Please extract the technological innovation points", etc.; for market - type texts, the templates may contain "Please analyze the market positioning", "Please extract the competitive advantages", etc. These templates are stored in the prompt template library and establish a mapping relationship with the corresponding types of business text, forming a dynamically updatable template system.
[0041] S102. Extract features from the business text data through the prompt templates to obtain initial feature data.
[0042] After obtaining the corresponding prompt templates, it is necessary to extract features from the enterprise business text data to obtain initial feature data. Initial feature data refers to the structured information extracted from the original business text, and this information will be used as the basic data for constructing the enterprise portrait. The reason for feature extraction is that the original business text data is usually unstructured natural - language text, and it needs to be converted into a structured data format that can be processed by a computer.
[0043] In this embodiment, the prompt templates contain feature extraction rules, and the feature extraction rules mainly include text classification rules and entity recognition rules. The text classification rules are used to determine the category attributes of the text, such as classifying a patent document as R & D - type, application - type, etc.; the entity recognition rules are used to identify and extract specific types of entity information from the text, such as company names, technical terms, time and location, etc.
[0044] The specific feature extraction process is as follows: First, the system classifies the business text data according to the text classification rules to obtain classification features. For example, for a patent document of an enterprise, the system analyzes information such as the technical field description and application scenario in the document and classifies it into the classification feature of "Artificial Intelligence - Image Recognition - Application Category". This classification feature can reflect the enterprise's layout in different technical fields.
[0045] Next, the system uses entity recognition rules to perform entity recognition on the business text data to obtain entity features. For example, specific algorithm names, technical parameters, application scenarios and other entity information are identified from the patent document. These entity features can describe the enterprise's technical capabilities and innovation focuses in detail.
[0046] For the convenience of subsequent processing, the extracted features need to be vectorized. The system first vectorizes the classification features to obtain classification feature vectors. For example, a classification label such as "Artificial Intelligence - Image Recognition - Application Category" is converted into a multi-dimensional numerical vector. Similarly, the system also vectorizes the entity features to obtain entity feature vectors. In the vectorization process, the system uses a pre-trained word embedding model to convert text features into numerical vectors of a fixed dimension. These vectors can provide a computable mathematical representation while preserving the original semantic information.
[0047] Finally, the system concatenates the classification feature vectors and entity feature vectors to obtain the complete initial feature data. The concatenation operation integrates feature information of different dimensions into a unified feature space to form a comprehensive description of the enterprise's business text. For example, for an image recognition patent document of a technology company, the initial feature data obtained after feature extraction contains both the vector representation of technical classification information and the vector representation of specific entities such as algorithms and application scenarios. These features together constitute the digital expression of the patent document.
[0048] Based on the above embodiments, as an optional implementation manner, in S012, feature extraction is performed on the business text data through a prompt template to obtain the initial feature data, which specifically includes S21 - S25: S21, obtain the feature extraction rules in the prompt template, and the feature extraction rules include text classification rules and entity recognition rules.
[0049] The feature extraction rules mainly include two important components: text classification rules and entity recognition rules. The text classification rules are used to determine the category attributes of the text, while the entity recognition rules are used to identify specific entity information in the text. This dual rule system can comprehensively analyze the text from both macroscopic and microscopic levels.
[0050] S22. Classify the business text data according to the text classification rules to obtain classification features.
[0051] S23. Identify entities in the business text data according to the entity recognition rules to obtain entity features.
[0052] S24. Vectorize the classification features to obtain classification feature vectors; vectorize the entity features to obtain entity feature vectors.
[0053] After obtaining the feature extraction rules, the system first applies the text classification rules to classify the business text data. The classification process adopts a hierarchical classification strategy, that is, first perform coarse-grained classification and then fine-grained classification. For example, for a technical document, the system first classifies it as a "technical document", and then further divides it into "R & D category" or "application category", and finally may be specific to a fine category such as "deep learning algorithm - computer vision - object detection". The system uses a pre-trained classification model in the classification process. This model is trained based on large-scale labeled data and can accurately identify the category features of the text. Through this classification process, the system obtains classification features that reflect the text theme and attributes.
[0054] Next, the system uses the entity recognition rules to identify entities in the business text data. The entity recognition process adopts a method that combines a deep learning model with rule constraints and can accurately identify key entity information in the text. For example, when processing an AI technology patent document, the system can identify specific entities such as algorithm names (such as "YOLO", "ResNet"), performance indicators (such as "accuracy 99.5%"), and application scenarios (such as "intelligent manufacturing quality inspection"). These entity information constitute entity features, which detail the specific technical content and key information points in the text.
[0055] To enable the extracted features to be efficiently processed by the computer, it is necessary to vectorize the classification features and entity features. The vectorization process uses a pre-trained language model, such as BERT or GPT, etc., to convert the text features into numerical vectors of a fixed dimension. For classification features, the system uses a multi-label encoding method to convert the hierarchical classification information into a representation in a high-dimensional vector space. For example, a classification feature such as "deep learning algorithm - computer vision - object detection" may be converted into a 300-dimensional vector. For entity features, the system uses an entity embedding model to encode the identified entities and their attribute information into numerical vectors. In this way, each entity feature is converted into a vector representation with semantic information.
[0056] S25. Concatenate the classification feature vectors and entity feature vectors to obtain initial feature data.
[0057] Finally, the system needs to concatenate the categorical feature vector and the entity feature vector to obtain the complete initial feature data. The concatenation process takes into account the importance weights of the two types of features and uses a weighted concatenation method. For example, if the categorical feature vector is 300-dimensional and the entity feature vector is 500-dimensional, the concatenated vector may be an 800-dimensional feature vector, where the weights of different dimensions are dynamically adjusted according to the importance of the features.
[0058] S103, input the initial feature data into a preset model to generate an initial enterprise portrait of the enterprise, and correct the preset model through the initial enterprise portrait to obtain a target model.
[0059] After obtaining the initial feature data, it is necessary to generate an initial enterprise portrait of the enterprise through a preset model. The preset model is a pre-trained deep learning model used to convert the feature data of the enterprise into a structured enterprise portrait representation. The reason for using the preset model is that the business characteristics of the enterprise are often multi-dimensional and multi-level, and a powerful model is needed to understand and integrate these complex feature information. The initial enterprise portrait refers to the enterprise feature description first generated by the model, which includes preliminary characterizations of the enterprise in multiple dimensions such as technical capabilities, business directions, and market positions.
[0060] In this embodiment, the training process of the preset model is completed on a large-scale labeled sample data set. The labeled sample data set contains the feature data of a large number of known enterprises and the corresponding portrait labels, and these data are divided into a training sample set and a validation sample set according to a preset ratio of 8:2. The system first constructs an initial preset model based on the training sample set, and then inputs the validation sample set into the model for validation to obtain a validation result. The validation result includes performance indicators such as the accuracy and recall rate of the model. If the validation result fails to meet the preset conditions (for example, the accuracy is lower than 90%), the system will optimize the initial preset model by adjusting model parameters, optimizing the network structure, etc. until the validation result meets the requirements, and finally obtain a preset model with good performance.
[0061] After inputting the initial feature data into the preset model, the model will generate an initial enterprise portrait of the enterprise. For example, for an enterprise focusing on computer vision technology, the initial enterprise portrait may include feature descriptions such as "core technology field: image recognition", "technology maturity: high", "market application fields: security monitoring, industrial quality inspection". However, since the preset model is trained based on general data, it may not fully adapt to the unique characteristics of a specific enterprise. Therefore, it is necessary to correct the preset model through the initial enterprise portrait to improve the model's ability to understand the characteristics of the enterprise.
[0062] The correction process first extracts feature parameters from the initial enterprise portrait. These parameters reflect the specific feature information of the enterprise and are used as training data. Based on this training data, the system iteratively adjusts the model parameters of the preset model, and a set of candidate model parameters is generated in each iteration. To evaluate the effect of parameter adjustment, the system calculates the performance evaluation values of each candidate model parameter. The performance evaluation value is calculated through predefined evaluation metrics (such as classification accuracy, feature extraction accuracy, etc.) and reflects the fitting degree of the model to the enterprise features. The system selects the candidate model parameter with the largest performance evaluation value as the target parameter and updates the model parameters of the preset model with this set of target parameters to finally obtain a target model with stronger adaptability.
[0063] Based on the above embodiments, as an alternative implementation manner, before inputting the initial feature data into the preset model to generate the initial enterprise portrait in S103, it specifically further includes S31 - S34: S31, obtain the labeled sample data set, and divide the labeled sample data set into a training sample set and a validation sample set according to a preset ratio.
[0064] To build a high-performance preset model, it is necessary to perform model training and optimization based on a large amount of labeled sample data. First, obtain the labeled sample data set, which contains the feature data of a large number of known enterprises and the corresponding enterprise portrait labels. The sources of the labeled sample data include existing enterprise portrait case libraries, manually labeled enterprise data, etc. To ensure the generalization ability of the model, the system divides the labeled sample data set into a training sample set and a validation sample set according to a preset ratio of 8:2. This division ratio is the optimal ratio verified through a large number of experiments, which can not only ensure the sufficiency of the training data but also guarantee the effectiveness of the validation.
[0065] S32, build an initial preset model based on the training sample set.
[0066] The process of building the initial preset model based on the training sample set adopts a deep neural network architecture. Specifically, the initial preset model is a multi-layer neural network structure, which includes, from bottom to top: a feature input layer, a feature extraction layer, a feature fusion layer, a portrait generation layer, and an output layer. The feature input layer is responsible for receiving the initial feature data, adopts a fully connected network structure, and the input dimension matches the dimension of the feature vector. The feature extraction layer consists of multiple convolutional neural network blocks, and each convolutional block contains a convolutional layer, a batch normalization layer, and an activation function layer, which are used to extract the local and global information of the features. The feature fusion layer combines the attention mechanism and the residual connection, which can effectively fuse the feature information at different levels. The portrait generation layer uses a multi-head self-attention mechanism, which can capture the complex associations between features. The output layer maps the features to a standardized enterprise portrait representation through the softmax function.
[0067] The training process of the model adopts an iterative optimization strategy. First, the loss function is defined, which includes two parts: classification loss and feature reconstruction loss. The classification loss uses the cross-entropy loss function to measure the difference between the predicted portrait and the true annotation; the feature reconstruction loss uses the mean squared error loss function to ensure that the model can retain the key information of the original features. The Adam optimizer is selected as the optimizer, and the cosine annealing strategy is adopted for the learning rate, with the initial learning rate set to 0.001. The batch training method is used during the training process, with each batch containing 64 samples, and a total of 100 epochs are trained. To prevent overfitting, techniques such as dropout and L2 regularization are also introduced.
[0068] S33. Input the validation sample set into the initial preset model for validation to obtain the validation result.
[0069] S34. Optimize the initial preset model according to the validation result until the validation result meets the preset conditions to obtain the preset model.
[0070] When inputting the validation sample set into the initial preset model for validation, the system will calculate multiple validation metrics. These metrics include the accuracy of the model (the matching degree between the predicted portrait and the true portrait), recall rate (the ability of the model to capture key features), F1 score (the harmonic mean of accuracy and recall rate), and feature fidelity (the degree to which the model retains the original feature information). For example, for the validation samples of a technology enterprise, the system will verify whether the model can accurately identify key features such as its technology field and innovation ability.
[0071] The process of optimizing the initial preset model according to the validation result is iterative. If the validation result fails to meet the preset conditions (for example, the accuracy is lower than 90% and the F1 score is lower than 0.85), the system will take a series of optimization measures. First is the network structure optimization, including adjusting the number of network layers, changing the convolution kernel size, modifying the attention mechanism parameters, etc. Second is the hyperparameter optimization, using methods such as grid search or Bayesian optimization to find the optimal learning rate, batch size, regularization parameter, etc. Finally is the ensemble learning optimization, which improves the model performance through model ensemble.
[0072] Based on the above embodiments, as an alternative implementation, in S103, correcting the preset model through the initial enterprise portrait to obtain the target model specifically includes S41 - S44: S41. Obtain the feature parameters in the initial enterprise portrait and use the feature parameters as training data.
[0073] To enable the preset model to better adapt to the characteristics of a specific enterprise, it is necessary to use the information in the initial enterprise portrait to make targeted corrections to the model. Taking a technology company specializing in autonomous driving technology as an example, the initial enterprise portrait of this company contains 200-dimensional feature parameters, which reflect the quantitative indicators of the enterprise in multiple dimensions such as technological innovation, market performance, and development potential. For example, the technological innovation dimension includes parameters such as "algorithm innovation ability: 0.95" and "patent quality index: 0.88", and the market performance dimension includes parameters such as "market share growth rate: 0.75" and "customer satisfaction: 0.92". These feature parameters are used as training data to guide the parameter adjustment of the preset model.
[0074] S42, iteratively adjust the model parameters of the preset model based on the training data to obtain multiple sets of candidate model parameters.
[0075] The iterative adjustment of the model parameters adopts a combination of gradient descent and genetic algorithm. First, the system uses the stochastic gradient descent method to fine-tune the parameters of different layers of the preset model, and at the same time introduces the genetic algorithm to explore a larger parameter space. In each iteration, the system will generate multiple sets of candidate model parameters. Taking this autonomous driving enterprise as an example, the system may try different weight configurations in the network layer for processing technology-related features in view of its high-tech innovation characteristics. For example, it may generate a set of parameter configurations that strengthen the weights of technological innovation (weight coefficient 1.2), a set of parameter configurations that balance technology and the market (weight coefficient 1.0), etc., to form 10 different candidate model parameter groups.
[0076] S43, calculate the performance evaluation values of each candidate model parameter, and take the candidate model parameter with the largest performance evaluation value as the target parameter.
[0077] S44, update the model parameters of the preset model to the target parameters to obtain the target model.
[0078] For each set of candidate model parameters, the system calculates its performance evaluation value through comprehensive evaluation indicators. The evaluation indicators include model prediction accuracy (accounting for 40%), feature fidelity (accounting for 30%), and generalization ability (accounting for 30%). Taking a certain candidate parameter group as an example, when processing the data of this autonomous driving enterprise, if its prediction accuracy reaches 96%, feature fidelity reaches 94%, and the generalization ability score is 0.91, then its comprehensive performance evaluation value is 0.94. By comparing the performance evaluation values of all candidate parameter groups, the system selects the parameter group with the highest evaluation value as the target parameter.
[0079] Finally, the system updates the parameters of the preset model to the target parameters, completing the customized correction of the model. This update process adopts a smooth transition strategy, that is, by introducing a decay factor (such as 0.8) to retain part of the generality of the original model, while integrating new target parameters to enhance the sensitivity of the model to specific enterprise characteristics. After such correction, the obtained target model shows better performance when processing the portrait generation tasks of this autonomous driving enterprise and similar enterprises.
[0080] S104, retrieve multiple external business information of the enterprise, calculate the relevance values between each external business information and the initial enterprise portrait, and input the external business information with relevance values greater than the preset threshold into the target model to obtain supplementary features.
[0081] Since the internal business text data of the enterprise may have incomplete information or be updated in a timely manner, the initial enterprise portrait generated only relying on internal data may not comprehensively reflect the actual situation of the enterprise. Therefore, it is necessary to supplement and improve the enterprise portrait by retrieving the external business information of the enterprise. External business information includes but is not limited to public information such as industry analysis reports, news, market research reports, competitor analysis, and expert comments.
[0082] In this embodiment, the system first retrieves external business information related to the enterprise through public data sources. The retrieval process adopts a multi-source retrieval strategy, that is, obtaining information from multiple data sources simultaneously. For example, for an enterprise focusing on artificial intelligence technology, the system will retrieve information such as technical discussions of the enterprise on professional technical forums, reports of industry media, and analysis reports of third-party consulting agencies. To ensure the timeliness of the retrieval results, the system will preferentially obtain information within the most recent year and conduct a preliminary screening according to the authority of the information source.
[0083] After obtaining the external business information, it is necessary to calculate the relevance values between this information and the initial enterprise portrait to screen out the most valuable supplementary information. The relevance value is a numerical index measuring the matching degree between external information and the existing characteristics of the enterprise, with a value range of 0 to 1, and the larger the value, the higher the relevance. The specific calculation method is to convert both the external business information and the initial enterprise portrait into feature vectors, and then calculate the similarity degree between the vectors through algorithms such as cosine similarity. For example, if an industry report details the technological innovation of the enterprise in the field of computer vision, and the initial enterprise portrait also contains relevant technical characteristics, then the relevance value of this report may be relatively high.
[0084] The system compares the relevance value with a preset threshold to filter out highly relevant external business information. The preset threshold is usually set to 0.7 or higher, which is determined through a large amount of experimental data and can effectively filter out weakly relevant information. For external business information with a relevance value greater than the preset threshold, the system performs text preprocessing on it, including operations such as removing noise information and unifying formats, to obtain preprocessed text.
[0085] The preprocessed text is then input into the feature extraction layer of the target model. The feature extraction layer is a network layer in the target model specifically responsible for extracting text features. It can extract key semantic features from the text and generate feature vectors. For example, extract vector representations of key information such as a company's technological advantages, market positioning, and development strategies from an analysis report.
[0086] To integrate the features of multiple external information sources, the system performs feature fusion on the extracted feature vectors based on preset fusion rules. The fusion rules include methods such as weighted average and attention mechanism, which are used to organically combine feature information from different sources to obtain fused features. Finally, the fused features are input into the feature mapping layer of the target model, and through the mapping and transformation of the deep neural network, the final supplementary features are obtained.
[0087] Based on the above embodiments, as an alternative implementation, in S104, inputting the external business information with a relevance value greater than the preset threshold into the target model to obtain supplementary features specifically includes S51 - S54: S51, perform text preprocessing on the external business information with a relevance value greater than the preset threshold to obtain preprocessed text.
[0088] First, perform text preprocessing on this highly relevant external business information. The preprocessing process includes text cleaning (removing special characters and unifying formats), word segmentation (identifying technical terms using a professional dictionary), stop word removal, etc. For example, for a technical blog of this AI chip company, the system identifies key technical indicators such as "7nm process", "300% increase in computing power", and "50% reduction in power consumption", and converts the text into a standardized format. This preprocessing process ensures the accuracy of subsequent feature extraction.
[0089] S52, extract the feature extraction layer from the target model, and input the preprocessed text into the feature extraction layer to obtain feature vectors.
[0090] Next, extract the feature extraction layer from the target model. The feature extraction layer of the target model is a sub-network containing multiple convolutional neural networks, which is specifically used to extract deep features from the text. Input the preprocessed text into the feature extraction layer, and the system will generate corresponding feature vectors. Taking this technology blog as an example, the feature extraction layer can capture features in multiple dimensions such as technological level, performance improvement, and energy consumption optimization, and encode them into a 500-dimensional feature vector. Each dimension in this feature vector corresponds to specific semantic information.
[0091] S53, perform feature fusion on the feature vectors based on a preset fusion rule to obtain fused features.
[0092] S54, input the fused features into the feature mapping layer of the target model to obtain supplementary features.
[0093] Then perform feature fusion on the feature vectors based on a preset fusion rule. The fusion rule combines the attention mechanism and weighted average, taking into account the credibility and importance of information from different sources. For example, for features from official technology blogs, the system assigns a higher weight (such as 0.8); for features from third-party analysis reports, a relatively lower weight (such as 0.6) is assigned. Through this weighted fusion method, the system integrates multiple feature vectors into a 300-dimensional fused feature. This fusion process ensures that important information is fully retained while filtering out redundant and noisy information.
[0094] Finally, input the fused features into the feature mapping layer of the target model to generate the final supplementary features. The feature mapping layer is a fully connected neural network responsible for mapping the fused features to the same feature space as the enterprise portrait. For this AI chip enterprise, the feature mapping layer converts the 300-dimensional fused features into 200-dimensional supplementary features. These supplementary features correspond to the feature dimensions of the enterprise portrait and contain quantitative indicators in multiple aspects such as technological innovation ability, market competitiveness, and development potential.
[0095] S105, input the initial enterprise portrait and supplementary features into the target model to generate the enterprise portrait label of the enterprise.
[0096] After obtaining the initial enterprise portrait and supplementary features, it is necessary to integrate this information and generate the final enterprise portrait label. The enterprise portrait label is a highly generalized and standardized expression of enterprise features, containing labeled descriptions of the enterprise in multiple dimensions such as technological capabilities, business directions, and development stages. The purpose of generating a standardized enterprise portrait label is to facilitate the rapid identification, precise matching, and efficient application of enterprise features, for example, playing an important role in scenarios such as business recommendation, risk assessment, and partner screening.
[0097] In this embodiment, first, it is necessary to merge the features of the initial enterprise portrait and the supplementary features. The initial enterprise portrait contains the basic features extracted from the enterprise's internal data, while the supplementary features contain the extended features obtained from external information. These two types of features have their own characteristics and complement each other. The feature merging process uses a feature fusion algorithm, which considers the importance weights of different features and organically integrates the two types of features through weighted combination. For example, for an artificial intelligence enterprise, its initial enterprise portrait may focus on the description of technical capabilities, while the supplementary features may include more market feedback and industry evaluations. Through feature merging, a complete feature representation that includes both the technical dimension and the market dimension can be obtained.
[0098] To ensure the comparability of features and the stability of calculations, the system normalizes the merged features to obtain standardized features. The normalization process includes the unification of the numerical range, the elimination of dimensions, the handling of outliers, etc. This can map features of different dimensions into the same numerical interval, facilitating subsequent model processing. For example, all numerical features are uniformly mapped to the interval [0, 1] to ensure fair comparison and calculation of features in different dimensions.
[0099] The standardized features are then input into the classification layer of the target model. The classification layer is a neural network layer that has been specifically trained and can map the input features into a set of predicted probability values. These predicted probability values reflect the matching degree of the enterprise in different label categories. For example, in the dimension of technological innovation ability, the model may output a probability of 0.85 for "leading type", a probability of 0.12 for "following type", and a probability of 0.03 for "starting type", indicating that the enterprise is most likely a technology-leading enterprise.
[0100] Finally, the system maps the predicted probability values into specific enterprise portrait labels according to the preset label mapping rules. The label mapping rules are a set of predefined conversion rules used to convert probability values into standardized label expressions. These rules not only consider the highest probability value in a single dimension but also the logical relationships between multiple dimensions to ensure that the generated label combinations are meaningful. For example, a certain technology enterprise may finally be labeled with a multi-dimensional label combination such as "technological innovation type - market leading level - rapid growth stage - high R & D investment - Series B financing stage".
[0101] Taking an AI chip enterprise as an example, by integrating its technical patent data (initial enterprise portrait) and the latest market feedback information (supplementary features), the system successfully generated accurate portrait labels such as "hard technology innovation enterprise - leader in the AI chip field - commercialization stage - domestic substitution direction". These labels not only accurately reflect the core features of the enterprise but also reflect the enterprise's positioning and development trend in the industry.
[0102] Based on the above embodiments, as an alternative implementation, in S105, inputting the initial enterprise portrait and supplementary features into the target model to generate enterprise portrait tags for the enterprise specifically includes S61 - S64: S61, merge the features of the initial enterprise portrait and supplementary features to obtain merged features.
[0103] To obtain accurate and easily understandable enterprise portrait tags, it is necessary to effectively integrate and process the initial enterprise portrait and supplementary features. Taking a high-tech enterprise focusing on quantum computing as an example, its initial enterprise portrait is a 200-dimensional feature vector, containing quantitative indicators such as basic technical capabilities, innovation levels, and market positions; the supplementary features are a 200-dimensional vector, reflecting dynamic information such as the enterprise's latest technological breakthroughs and market performance. The system combines feature splicing and weighted superposition to merge the features and obtains a 400-dimensional merged feature. During the merging process, for overlapping feature dimensions (such as technological innovation capabilities), the system assigns different weights according to timeliness. For example, a weight of 0.4 is assigned to the historical data in the initial enterprise portrait, and a weight of 0.6 is assigned to the latest data in the supplementary features, ensuring that the merged features retain both historical stability and reflect the latest changes.
[0104] S62, perform normalization processing on the merged features to obtain normalized features.
[0105] To eliminate the dimensional differences between different feature dimensions, the system performs normalization processing on the merged features. The normalization uses an improved Z-score normalization method, which not only considers the mean and variance of the features but also introduces an industry background factor to adjust the normalization process. For example, for the technical indicators of this quantum computing enterprise, the system will adjust the normalization parameters according to the overall development level of the quantum computing field. Specifically, if the original value of a certain technical indicator is 0.85, the industry mean is 0.6, the standard deviation is 0.15, and the industry background factor is 1.2, then the normalized value is [(0.85 - 0.6) / (0.15 * 1.2)]. After such processing, the normalized features can more accurately reflect the enterprise's relative position in the industry.
[0106] S63, input the normalized features into the classification layer of the target model to obtain predicted probability values.
[0107] When the standardized features are input into the classification layer of the target model, the system adopts a multi-layer softmax classifier structure, with each layer corresponding to portrait labels of different dimensions. The classification layer processes the standardized features through a deep neural network and outputs predicted probability values in multiple dimensions. Taking this quantum computing enterprise as an example, the classification layer will output probability distributions in multiple dimensions such as the type of technological innovation (e.g., the probability of "original technology type" is 0.92, and the probability of "technology application type" is 0.05), the development stage (e.g., the probability of "rapid growth stage" is 0.88, and the probability of "maturity stage" is 0.10), and the core competitiveness (e.g., the probability of "technology leading type" is 0.95, and the probability of "market-driven type" is 0.03).
[0108] S64, according to the preset label mapping rules, map the predicted probability values to the enterprise portrait labels of the corresponding enterprise.
[0109] Finally, the system maps these predicted probability values to specific enterprise portrait labels according to the preset label mapping rules. The label mapping rules not only consider the maximum probability value of a single dimension but also consider the logical relationships between multiple dimensions. For example, when the type of technological innovation is "original technology type" (probability 0.92) and the core competitiveness is "technology leading type" (probability 0.95), the system will label the enterprise as a "technology innovation leading enterprise"; when these labels are combined with the enterprise's development stage of "rapid growth stage" (probability 0.88), the final generated portrait label combination is "rapidly growing technology innovation leading enterprise".
[0110] Based on the above method, this application also discloses an enterprise user portrait understanding system, as Figure 2 shown Figure 2 is a schematic structural diagram of an enterprise user portrait understanding system provided by an embodiment of this application. The system includes: an acquisition module, an extraction module, a first input module, a calculation module, and a second input module; wherein, The acquisition module is used to acquire enterprise business text data and select a prompt word template corresponding to the type of the business text data from the prompt word template library; the extraction module is used to extract features from the business text data through the prompt word template to obtain initial feature data; the first input module is used to input the initial feature data into a preset model to generate an initial enterprise portrait of the enterprise, and correct the preset model through the initial enterprise portrait to obtain a target model; the calculation module is used to retrieve multiple external business information of the enterprise, calculate the correlation values between each external business information and the initial enterprise portrait, and input the external business information with a correlation value greater than a preset threshold into the target model to obtain supplementary features; the second input module is used to input the initial enterprise portrait and the supplementary features into the target model to generate the enterprise portrait labels of the enterprise.
[0111] It should be noted that when the system provided in the above embodiments realizes its functions, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.
[0112] Please refer to Figure 3 , which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 3 shown, the electronic device 1000 may include: at least one processor 1001, at least one network interface 1004, a user interface 1003, a memory 1005, and at least one communication bus 1002.
[0113] Among them, the communication bus 1002 is used to realize the connection and communication between these components.
[0114] Among them, the user interface 1003 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface.
[0115] Among them, the network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0116] Among them, the processor 1001 may include one or more processing cores. The processor 1001 connects various parts within the entire server through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 1005, and by calling the data stored in the memory 1005, it executes various functions of the server and processes data. Optionally, the processor 1001 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 1001 may integrate one or a combination of several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 1001 and may be implemented separately by a single chip.
[0117] Among them, the memory 1005 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory 1005 includes a non-transitory computer-readable storage medium. The memory 1005 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 1005 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area can store the data involved in the above-mentioned various method embodiments. Optionally, the memory 1005 may also be at least one storage device located far from the aforementioned processor 1001. As Figure 3 shown, the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for an enterprise user portrait understanding method.
[0118] In Figure 3In the electronic device 1000 shown, the user interface 1003 is mainly used to provide an interface for the user to input and obtain the data input by the user; and the processor 1001 can be used to call the application program stored in the memory 1005 for a method of understanding enterprise user portraits. When executed by one or more processors, the electronic device is caused to execute one or more of the methods as described in the above embodiments.
[0119] An electronic device-readable storage medium stores instructions. When executed by one or more processors, the electronic device is caused to execute one or more of the methods as described in the above embodiments.
[0120] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0121] In the above embodiments, the descriptions of the various embodiments each have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0122] In several embodiments provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.
[0123] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0124] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0125] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned memory includes: various media such as USB flash drives, mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0126] The foregoing are only exemplary embodiments of the present disclosure and should not be used to limit the scope of the present disclosure. That is, all equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. Those skilled in the art will readily think of other implementation manners of the present disclosure after considering the specification and practicing the disclosure herein. This application aims to cover any variations, uses, or adaptive changes of the present disclosure, and these variations, uses, or adaptive changes follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.
Claims
1. A method for understanding enterprise user portraits, characterized in that The method includes: Obtain enterprise business text data, and select a prompt word template corresponding to the type of the business text data from a prompt word template library; Extract features from the business text data through the prompt word template to obtain initial feature data; Input the initial feature data into a preset model to generate an initial enterprise portrait of the enterprise, and correct the preset model through the initial enterprise portrait to obtain a target model; Retrieve multiple external business information of the enterprise, calculate the correlation values between each external business information and the initial enterprise portrait, and input the external business information with the correlation value greater than a preset threshold into the target model to obtain supplementary features; Input the initial enterprise portrait and the supplementary features into the target model to generate an enterprise portrait label of the enterprise.
2. The enterprise user portrait understanding method according to claim 1, wherein Before selecting a prompt word template corresponding to the type of the business text data from the prompt word template library, it further includes: Obtain a business text data set of the enterprise, classify the business text data set to obtain multiple data sets; Select keywords corresponding to the text features of each data set from a preset keyword library; Label each data set according to the keywords to obtain various types of business text data; Generate corresponding prompt word templates according to each type of business text data, store the prompt word templates in the prompt word template library, and establish a mapping relationship between each type of business text data and the corresponding prompt word template.
3. The method for understanding enterprise user portraits according to claim 1, wherein The extracting features from the business text data through the prompt word template to obtain initial feature data includes: Obtain the feature extraction rules in the prompt word template, where the feature extraction rules include text classification rules and entity recognition rules; Classify the business text data according to the text classification rules to obtain classification features; Perform entity recognition on the business text data according to the entity recognition rules to obtain entity features; Perform vectorization processing on the classification features to obtain classification feature vectors; perform vectorization processing on the entity features to obtain entity feature vectors; Concatenate the classification feature vectors and the entity feature vectors to obtain initial feature data.
4. The method for understanding enterprise user portraits according to claim 1, wherein Before inputting the initial feature data into a preset model to generate an initial enterprise portrait of the enterprise, it further includes: Obtain a labeled sample data set, and divide the labeled sample data set into a training sample set and a validation sample set according to a preset ratio; Construct an initial preset model based on the training sample set; Input the validation sample set into the initial preset model for validation to obtain a validation result; Optimize the initial preset model according to the validation result until the validation result meets the preset conditions to obtain a preset model.
5. The method for understanding enterprise user portraits according to claim 1, wherein The correcting the preset model through the initial enterprise portrait to obtain a target model includes: Obtain the feature parameters in the initial enterprise portrait, and use the feature parameters as training data; Iteratively adjust the model parameters of the preset model based on the training data to obtain multiple groups of candidate model parameters; Calculate the performance evaluation values of each of the candidate model parameters, and use the candidate model parameter with the largest performance evaluation value as the target parameter; Update the model parameters of the preset model to the target parameters to obtain a target model.
6. The method for understanding enterprise user portraits according to claim 1, wherein The step of inputting the external business information with a relevance value greater than a preset threshold into the target model to obtain supplementary features includes: Perform text preprocessing on the external business information with a relevance value greater than a preset threshold to obtain preprocessed text; Extract a feature extraction layer from the target model, and input the preprocessed text into the feature extraction layer to obtain feature vectors; Perform feature fusion on the feature vectors based on a preset fusion rule to obtain fused features; Input the fused features into the feature mapping layer of the target model to obtain supplementary features.
7. The method for understanding enterprise user portraits according to claim 1, wherein The step of inputting the initial enterprise portrait and the supplementary features into the target model to generate the enterprise portrait label of the enterprise includes: Merge the features of the initial enterprise portrait and the supplementary features to obtain merged features; Perform normalization processing on the merged features to obtain normalized features; Input the normalized features into the classification layer of the target model to obtain predicted probability values; According to a preset label mapping rule, map the predicted probability values to the enterprise portrait labels corresponding to the enterprise.
8. An enterprise user portrait understanding system, characterized in that, The system includes: an acquisition module, an extraction module, a first input module, a calculation module, and a second input module; wherein, The acquisition module is configured to acquire enterprise business text data and select a prompt word template corresponding to the type of the business text data from a prompt word template library; The extraction module is configured to extract features from the business text data through the prompt word template to obtain initial feature data; The first input module is configured to input the initial feature data into a preset model to generate an initial enterprise portrait of the enterprise, and correct the preset model through the initial enterprise portrait to obtain a target model; The calculation module is configured to retrieve multiple external business information of the enterprise, calculate the relevance values between each external business information and the initial enterprise portrait, and input the external business information with a relevance value greater than a preset threshold into the target model to obtain supplementary features; The second input module is configured to input the initial enterprise portrait and the supplementary features into the target model to generate the enterprise portrait label of the enterprise.
9. An electronic device, characterized in that, It includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. The user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory so that the electronic device executes the method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, A computer program is stored that can be loaded and executed by a processor to execute the method according to any one of claims 1-7.
Citation Information
Cited By
Object information acquisition method and device, storage medium and program product
CN120851220A
Object information acquisition method, device, storage medium, and program product
CN120851220B
Image generation method and device, equipment, storage medium and product
CN120997628A