Method, apparatus, device and storage medium for generating express delivery industry portrait

By performing word segmentation, feature extraction and cluster analysis on the original express data, combined with the prediction of the portrait generator, a multi-dimensional express industry portrait is built, solving the problem of portrait accuracy and inefficiency in the express industry, and achieving more efficient and accurate portrait generation.

CN112560474BActive Publication Date: 2025-06-20SHANGHAI DONGPU INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010944984.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-10
Publication Date
2025-06-20
Estimated Expiration
2040-09-10

AI Technical Summary

Technical Problem

The original express data in the express delivery industry is scattered and the data utilization rate is not high, resulting in the low accuracy and inefficiency of the express delivery industry portrait generated through the original express delivery data.

Method used

By obtaining the original express data, using a preset word segmenter to segment the segments, perform calculation processing to obtain processing data; using the feature extractor to extract feature vectors and feature labels; analyzing feature labels through clustering algorithms to generate multi-dimensional weight labels; inputting the weight labels into the portrait generator for prediction, and building a express industry portrait.

Benefits of technology

The accuracy and efficiency of generating express industry portraits through original express data has been improved, and the generated portraits are more accurate and multi-dimensional, which can better reflect the real situation in the express industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112560474B_ABST
    Figure CN112560474B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence, and discloses a method, device, equipment and storage medium for generating a portrait of the express delivery industry, which is used to improve the accuracy and efficiency of generating user portraits through user data. The method for generating a portrait of the express delivery industry includes: segmenting the text segments in the original express delivery data based on a preset word segmenter to obtain a word corpus, and obtaining processed data by performing calculation processing on the word corpus; using a preset feature extractor to extract features from the processed data to obtain feature vectors, and determining the feature labels corresponding to the processed data according to the feature vectors; classifying and analyzing the feature labels through a preset clustering algorithm to obtain multi-dimensional weight labels, and the weight labels at least include platform labels, address labels, time labels, commodity labels, user labels and merchant labels; using a preset portrait generator to predict the weight labels to obtain predicted labels, and constructing a portrait of the express delivery industry through the feature labels, weight labels and predicted labels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular, to a method, apparatus, device, and storage medium for generating a portrait of the express delivery industry. Background Art

[0002] With the rapid development of the economy, more and more users use online platforms to purchase goods, so the express delivery industry is developing more and more rapidly. Generally, there is a vast amount of user data in the express delivery industry. When processing data, the user data will generate a huge amount of data that is difficult to manage. However, with the popularization and promotion of cloud computing technology, it has made it possible to manage and intelligently analyze the real-time dynamics of a vast number of users. Taking the user portrait technology as an example, the user portrait technology analyzes a vast amount of data and then discovers the potential commercial value behind the data.

[0003] A user portrait analyzes and abstracts the complete information of a user by collecting and analyzing data such as the user's social attributes, living habits, and consumption behaviors. The user portrait system can summarize the shopping characteristics of users by analyzing their consumption habits and historical data, and can also obtain the demand information of users through the communication between sellers and users. With the buyer user portrait, it is helpful to implement precision marketing and precision positioning in commercial service promotion.

[0004] Since the original express delivery data in the express delivery industry is scattered and the data utilization rate is not high, the accuracy and efficiency of generating a portrait of the express delivery industry through the original express delivery data are not high. Summary of the Invention

[0005] The present invention provides a method, apparatus, device, and storage medium for generating a portrait of the express delivery industry, which is used to improve the accuracy and efficiency of generating a portrait of the express delivery industry through the original express delivery data.

[0006] In a first aspect of the present invention, a method for generating a portrait of the express delivery industry is provided, including: obtaining original express delivery data, segmenting the text segments in the original express delivery data based on a preset word segmenter to obtain word corpora, performing calculation processing on the word corpora to obtain processed data; using the preset feature extractor to extract features from the processed data to obtain a feature vector of the processed data, and determining a feature label corresponding to the processed data according to the feature vector; classifying and analyzing the feature labels through a preset clustering algorithm to obtain multi-dimensional weight labels, where the weight labels at least include platform labels, address labels, time labels, commodity labels, user labels, and merchant labels; inputting the weight labels into a preset portrait generator, using the preset portrait generator to predict the weight labels to obtain predicted labels, and constructing a portrait of the express delivery industry through the feature labels, the weight labels, and the predicted labels.

[0007] Optionally, in the first implementation manner of the first aspect of the present invention, the obtaining of the original express delivery data, segmenting the text segments in the original express delivery data based on a preset word segmenter to obtain word corpora, and obtaining processed data by performing calculation processing on the word corpora includes: obtaining the original express delivery data and transmitting the original express delivery data to the preset word segmenter; segmenting the text segments in the original express delivery data into multiple word corpora in the preset word segmenter, and counting the number of the multiple word corpora, where the word corpora are words or phrases existing in a standard dictionary; using a preset statistical function to count the frequency of occurrence of each word corpus in the original express delivery data to obtain multiple basic frequencies; calculating the number of times each word corpus appears in the text segment through each basic frequency to obtain multiple word frequencies, and calculating the inverse document frequency of each word corpus to obtain multiple inverse document frequencies, and determining multiple target word corpora according to the multiple word frequencies and the multiple inverse document frequencies to obtain processed data.

[0008] Optionally, in the second implementation manner of the first aspect of the present invention, the calculating the number of times each word corpus appears in the text segment through each basic frequency to obtain multiple word frequencies, and calculating the inverse document frequency of each word corpus to obtain multiple inverse document frequencies, and determining multiple target word corpora according to the multiple word frequencies and the multiple inverse document frequencies to obtain processed data includes: obtaining candidate corpora in the word corpora, calculating the number of times the candidate corpora appear in the text segment through the basic frequency corresponding to the candidate corpora and a preset first calculation formula to obtain target word frequencies, where the preset first calculation formula is:

[0009]

[0010] where TF is the target word frequency of the candidate corpus, n is the number of times the candidate corpus appears in the text segment, s is the number of all word corpora in the text segment, and both n and s are positive integers; calculating the inverse document frequency of the candidate corpus using a preset second calculation formula, where the preset second calculation formula is:

[0011]

[0012] Among them, IDF is the target inverse corpus frequency of the candidate corpus, q is the number of segments, z is the number of segments with candidate corpus, and both q and z are positive integers; obtain the remaining corpus in the word corpus except the candidate corpus, and calculate the remaining word frequency and the remaining inverse corpus frequency of the remaining corpus through the preset first calculation formula and the preset second calculation formula, merge the target word frequency and the remaining word frequency to obtain multiple word frequencies, and merge the target inverse corpus frequency and the remaining inverse corpus frequency to obtain multiple inverse corpus frequencies; screen out multiple target word corpora in multiple word corpora whose word frequency is greater than or equal to the first set threshold and whose inverse corpus frequency is less than or equal to the second set threshold, and determine the segments corresponding to the multiple target word corpora as the processed data.

[0013] Optionally, in the third implementation manner of the first aspect of the present invention, the using the preset feature extractor to extract features from the processed data to obtain the feature vector of the processed data, and determining the feature label corresponding to the processed data according to the feature vector includes: sending the processed data to the preset feature extractor, and using the preset feature extractor to extract features from the target word corpus in the processed data to obtain a feature vector; calculating the similarity between the feature vector and the label vector to obtain a basic similarity; selecting the target similarity with the largest numerical value of the basic similarity, and determining the preset label corresponding to the label vector for calculating the target similarity as the feature label corresponding to the processed data.

[0014] Optionally, in the fourth implementation manner of the first aspect of the present invention, the classifying and analyzing the feature labels through a preset clustering algorithm to obtain multi-dimensional weight labels, where the weight labels at least include platform labels, address labels, time labels, product labels, user labels, and merchant labels include:

[0015] Using a preset clustering function to select candidate labels from the feature labels; through a clustering algorithm, clustering the remaining labels with the candidate labels as the center to obtain grouped clustering labels, where the remaining labels are used to indicate the labels in the feature labels other than the candidate labels; extracting the keywords of the grouped clustering labels, and determining the keywords as the weight labels corresponding to the grouped clustering labels, where the keywords are the central words of the grouped clustering labels, and the weight labels at least include platform labels, address labels, time labels, product labels, user labels, and merchant labels.

[0016] Optionally, in the fifth implementation manner of the first aspect of the present invention, the steps of inputting the weight label into a preset portrait generator, using the preset portrait generator to predict the weight label to obtain a predicted label, and constructing an express delivery industry portrait through the feature label, the weight label, and the predicted label include: inputting the weight label into a preset portrait generator, using a preset logistic regression model in the preset portrait generator to predict the weight label to obtain a first predicted label; using a preset product diffusion model in the preset portrait generator to predict the weight label to obtain a second predicted label; using a preset churn warning model in the preset portrait generator to predict the weight label to obtain a third predicted label; merging the first predicted label, the second predicted label, and the third predicted label to obtain a predicted label; and inputting the feature label, the weight label, and the predicted label into a system construction model in the preset portrait generator to generate an express delivery industry portrait.

[0017] The second aspect of the present invention provides a device for generating an express delivery industry portrait, including: a processing module, configured to obtain original express delivery data, segment a text segment in the original express delivery data based on a preset tokenizer to obtain a word corpus, and obtain processed data by performing calculation processing on the word corpus; a determination module, configured to use a preset feature extractor to extract features of the processed data to obtain a feature vector of the processed data, and determine a feature label corresponding to the processed data according to the feature vector; a classification module, configured to perform classification analysis on the feature labels through a preset clustering algorithm to obtain multi-dimensional weight labels, where the weight labels at least include a platform label, an address label, a time label, a commodity label, a user label, and a merchant label; and a generation module, configured to input the weight labels into a preset portrait generator, use the preset portrait generator to predict the weight labels to obtain predicted labels, and construct an express delivery industry portrait through the feature labels, the weight labels, and the predicted labels.

[0018] Optionally, in the first implementation manner of the second aspect of the present invention, the processing module includes: an acquisition unit configured to acquire original express delivery data and transmit the original express delivery data to a preset word segmenter; a segmentation unit configured to segment the text segments in the original express delivery data into multiple word corpora in the preset word segmenter and count the number of the multiple word corpora, where the word corpus is a word or phrase existing in a standard dictionary; a statistics unit configured to use a preset statistical function to count the frequency of occurrence of each word corpus in the original express delivery data to obtain multiple basic frequencies; a determination unit configured to calculate the number of occurrences of the corresponding word corpus in the text segment through each basic frequency to obtain multiple word frequencies, and calculate the inverse document frequency of each word corpus to obtain multiple inverse document frequencies, and determine multiple target word corpora according to the multiple word frequencies and the multiple inverse document frequencies to obtain processed data.

[0019] Optionally, in the second implementation manner of the second aspect of the present invention, the determination unit is specifically configured to: acquire candidate corpora in the word corpus, calculate the number of occurrences of the candidate corpus in the text segment through the basic frequency corresponding to the candidate corpus and a preset first calculation formula to obtain a target word frequency, where the preset first calculation formula is:

[0020]

[0021] where TF is the target word frequency of the candidate corpus, n is the number of occurrences of the candidate corpus in the text segment, s is the number of all word corpora in the text segment, and both n and s are positive integers; calculate the inverse document frequency of the candidate corpus by using a preset second calculation formula, where the preset second calculation formula is:

[0022]

[0023] where IDF is the target inverse document frequency of the candidate corpus, q is the number of text segments, z is the number of text segments where the candidate corpus exists, and both q and z are positive integers; acquire the remaining corpora in the word corpus except the candidate corpus, calculate the remaining word frequency and the remaining inverse document frequency of the remaining corpus through the preset first calculation formula and the preset second calculation formula, merge the target word frequency and the remaining word frequency to obtain multiple word frequencies, and merge the target inverse document frequency and the remaining inverse document frequency to obtain multiple inverse document frequencies; screen out multiple target word corpora with a word frequency greater than or equal to a first set threshold and an inverse document frequency less than or equal to a second set threshold from the multiple word corpora, and determine the text segments corresponding to the multiple target word corpora as processed data.

[0024] Optionally, in the third implementation manner of the second aspect of the present invention, the determining module is specifically configured to: send the processed data to a preset feature extractor, use the preset feature extractor to extract features from the target word corpus in the processed data to obtain a feature vector; calculate the similarity between the feature vector and a label vector to obtain a basic similarity; select a target similarity with the largest value of the basic similarity, and determine the preset label corresponding to the label vector for calculating the target similarity as the feature label corresponding to the processed data.

[0025] Optionally, in the fourth implementation manner of the second aspect of the present invention, the classifying module is specifically configured to: select candidate labels from the feature labels by using a preset clustering function; perform clustering on the remaining labels with the candidate labels as the center through a clustering algorithm to obtain grouped clustering labels, where the remaining labels are used to indicate the labels other than the candidate labels in the feature labels; extract keywords of the grouped clustering labels, and determine the keywords as weight labels corresponding to the grouped clustering labels, where the keywords are the central words of the grouped clustering labels, and the weight labels at least include platform labels, address labels, time labels, commodity labels, user labels, and merchant labels.

[0026] Optionally, in the fifth implementation manner of the second aspect of the present invention, the generating module is specifically configured to: input the weight labels into a preset portrait generator, use a preset logistic regression model in the preset portrait generator to predict the weight labels to obtain a first predicted label; use a preset product diffusion model in the preset portrait generator to predict the weight labels to obtain a second predicted label; use a preset churn warning model in the preset portrait generator to predict the weight labels to obtain a third predicted label; merge the first predicted label, the second predicted label, and the third predicted label to obtain a predicted label; input the feature label, the weight label, and the predicted label into a system construction model in the preset portrait generator to generate an express delivery industry portrait.

[0027] The third aspect of the present invention provides a device for generating an express delivery industry portrait, including: a memory and at least one processor, where instructions are stored in the memory; the at least one processor calls the instructions in the memory so that the device for generating an express delivery industry portrait executes the above-mentioned method for generating an express delivery industry portrait.

[0028] The fourth aspect of the present invention provides a computer-readable storage medium, where instructions are stored in the computer-readable storage medium, and when the instructions run on a computer, the computer is enabled to execute the above-mentioned method for generating an express delivery industry portrait.

[0029] In the technical solution provided by the present invention, original express delivery data is obtained, and the paragraphs in the original express delivery data are segmented based on a preset word segmenter to obtain word corpora. By performing calculation processing on the word corpora, processed data is obtained; a preset feature extractor is used to extract features from the processed data to obtain a feature vector of the processed data, and a feature label corresponding to the processed data is determined according to the feature vector; a preset clustering algorithm is used to perform classification analysis on the feature labels to obtain multi-dimensional weight labels, and the weight labels at least include a platform label, an address label, a time label, a commodity label, a user label, and a merchant label; the weight labels are input into a preset portrait generator, and the preset portrait generator is used to predict the weight labels to obtain predicted labels, and an express delivery industry portrait is constructed through the feature labels, the weight labels, and the predicted labels. In the embodiments of the present invention, by inputting the original express delivery data into a preset word segmenter and a preset feature extractor for processing, feature labels corresponding to the original express delivery data are obtained, and then the preset clustering algorithm and the preset portrait generator are used to analyze and predict the feature labels to generate a multi-dimensional express delivery industry portrait, improving the accuracy and efficiency of generating an express delivery industry portrait using the original express delivery data. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 FIG. is a schematic diagram of an embodiment of a method for generating an express delivery industry portrait in an embodiment of the present invention;

[0031] Figure 2 FIG. is a schematic diagram of another embodiment of a method for generating an express delivery industry portrait in an embodiment of the present invention;

[0032] Figure 3 FIG. is a schematic diagram of an embodiment of an apparatus for generating an express delivery industry portrait in an embodiment of the present invention;

[0033] Figure 4 FIG. is a schematic diagram of another embodiment of an apparatus for generating an express delivery industry portrait in an embodiment of the present invention;

[0034] Figure 5 FIG. is a schematic diagram of an embodiment of a device for generating an express delivery industry portrait in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0035] Embodiments of the present invention provide a method, apparatus, device, and storage medium for generating an express delivery industry portrait, which are used to improve the accuracy and efficiency of generating an express delivery industry portrait using original express delivery data.

[0036] In the description, claims and the above drawings of the present invention, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order different from that shown or described herein. In addition, the term "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0037] For ease of understanding, the specific process of the embodiments of the present invention will be described below. Please refer to Figure 1 , an embodiment of the method for generating a portrait of the express delivery industry in the embodiments of the present invention includes:

[0038] 101. Obtain the original express delivery data, segment the text segments in the original express delivery data based on a preset word segmenter to obtain word corpora, and obtain processed data by performing calculation processing on the word corpora.

[0039] It can be understood that the execution subject of the present invention can be a device for generating a portrait of the express delivery industry, or a terminal or a server. Specifically, it is not limited here. The embodiments of the present invention will be described by taking the server as the execution subject as an example.

[0040] The server first needs to obtain the original express delivery data, input the original express delivery data into a preset word segmenter, perform segmentation processing on the text segments in the original express delivery data to obtain the word corpora in the text segments, and then the server performs calculations on the word corpora. By calculating the word frequency and inverse document frequency of the word corpora, processed data is obtained. Here, the original express delivery data can be the user's name, gender, contact information, delivery address, family information, commodity purchase date, commodity purchase record, commodity purchase attribute, etc. Input these original express delivery data into a preset word segmenter to process the original express delivery data.

[0041] 102. Use a preset feature extractor to extract features from the processed data to obtain a feature vector of the processed data, and determine a feature label corresponding to the processed data according to the feature vector.

[0042] The server transfers the processed data to a preset feature extractor, extracts features from the processed data through the preset feature extractor to obtain the feature vector of the processed data, and then determines the feature label corresponding to the processed data through the feature vector. The preset feature extractor here is used to extract features from the input data, and what is obtained after being processed by the preset feature extractor is the feature vector. The server then analyzes the feature vector to further determine the feature label of the original express delivery data.

[0043] 103. Classify and analyze the feature labels through a preset clustering algorithm to obtain multi-dimensional weight labels, where the weight labels at least include platform labels, address labels, time labels, commodity labels, user labels, and merchant labels;

[0044] The server classifies and analyzes the obtained feature labels through a preset clustering algorithm, and obtains representative weight labels through the classification and grouping of the feature labels.

[0045] The preset clustering algorithm here is the k-means clustering algorithm. The k-means clustering algorithm is an iterative clustering analysis algorithm. The algorithm principle is to randomly select K objects in the feature labels as the initial clustering centers, and then calculate the distance between each object and each seed clustering center, and assign each object to the clustering center closest to it. It should be noted that the clustering centers and the objects assigned to them represent a cluster. Each time a sample is assigned, the clustering center of the cluster will be recalculated based on the existing objects, and this process will continue to repeat until a certain termination condition is met. The server classifies and analyzes the feature labels through this algorithm to obtain multi-dimensional weight labels, where the weight labels at least include platform labels, address labels, time labels, commodity labels, user labels, and merchant labels.

[0046] 104. Input the weight labels into a preset portrait generator, use the preset portrait generator to predict the weight labels to obtain predicted labels, and construct an express delivery industry portrait through the feature labels, weight labels, and predicted labels.

[0047] The server predicts the weight labels to obtain predicted labels, and finally constructs an express delivery industry portrait in the preset portrait generator using the feature labels, weight labels, and predicted labels. The express delivery industry portrait refers to a labeled model abstracted from the information corresponding to multiple different dimensions of the express delivery industry. That is to say, labels are assigned to the express delivery industry from different dimensions, and the labels are highly refined feature identifiers obtained through the analysis of the original express delivery information. By assigning labels, some highly generalized and easily understood features can be used to describe the express delivery industry, making it easier for people to understand the express delivery industry and facilitating computer processing.

[0048] In an embodiment of the present invention, by inputting the original express delivery data into a preset word segmenter and a preset feature extractor for processing, feature tags corresponding to the original express delivery data are obtained. Then, through a preset clustering algorithm and a preset portrait generator, the feature tags are analyzed and predicted to generate a multi-dimensional portrait of the express delivery industry, improving the accuracy and efficiency of generating a portrait of the express delivery industry using the original express delivery data.

[0049] Please refer to Figure 2 , another embodiment of the method for generating a portrait of the express delivery industry in an embodiment of the present invention includes:

[0050] 201. Obtain the original express delivery data and transmit the original express delivery data to a preset word segmenter;

[0051] The server obtains the original express delivery data, inputs the original express delivery data into a preset word segmenter, and processes the original express delivery data to obtain processed data. Here, the original express delivery data can be the user's name, gender, contact information, delivery address, family information, commodity purchase date, commodity purchase record, commodity purchase attribute, etc. Input these original express delivery data into the preset word segmenter to process the original express delivery data.

[0052] 202. In the preset word segmenter, segment the paragraphs in the original express delivery data into multiple word corpora, and count the number of the multiple word corpora. The word corpora are words or phrases existing in a standard dictionary;

[0053] Since there are a large number of paragraphs or sentences in the original express delivery data, when processing it, it is first necessary to segment the paragraphs into multiple word corpora. It should be noted that when segmenting the paragraphs, the segmentation standard is to segment the paragraphs into words or phrases in a standard dictionary, that is, the paragraphs are composed of words or phrases in a standard dictionary. In addition, the process of segmenting the paragraphs into multiple word corpora here is a commonly used technical means in the art, so it will not be elaborated here. After segmenting the paragraphs into multiple word corpora, the server counts the number of all word corpora in the paragraphs. For example: the corresponding corpora segmented from "I am going to take a nap at school at noon today" are "today", "noon", "I", "prepare", "go", "school", "take a nap", rather than "this", "day noon", "I", "prepare", "go to learn", "school", "take a nap". The server counts the number of corresponding word corpora as 7.

[0054] 203. Use a preset statistical function to count the frequency of each word corpus appearing in the original express delivery data to obtain multiple basic frequencies;

[0055] For each word corpus, the server will use a preset statistical function to count the number of times the word corpus appears in the original express delivery data, and each word corpus has a basic frequency. It can be understood that if a word corpus appears multiple times in a text segment, it can be determined that the word corpus is a key word in the corresponding text segment, that is, the word corpus is the central word of the text segment.

[0056] 204. Calculate the number of times the corresponding word corpus appears in the text segment through each basic frequency to obtain multiple word frequencies, and calculate the inverse document frequency of each word corpus to obtain multiple inverse document frequencies. Determine multiple target word corpora based on the multiple word frequencies and multiple inverse document frequencies to obtain processed data;

[0057] Specifically, the server first obtains the candidate corpus in the word corpus, and calculates the number of times the candidate corpus appears in the text segment through the basic frequency corresponding to the candidate corpus and a preset first calculation formula to obtain the target word frequency. The preset first calculation formula is:

[0058]

[0059] Among them, TF is the target word frequency of the candidate corpus, n is the number of times the candidate corpus appears in the text segment, s is the number of all word corpora in the text segment, and both n and s are positive integers; secondly, the server uses a preset second calculation formula to calculate the inverse document frequency of the candidate corpus. The preset second calculation formula is:

[0060]

[0061] Among them, IDF is the target inverse document frequency of the candidate corpus, q is the number of text segments, z is the number of text segments containing the candidate corpus, and both q and z are positive integers; then the server obtains the remaining corpus in the word corpus except the candidate corpus, and calculates the remaining word frequency and remaining inverse document frequency of the remaining corpus through the preset first calculation formula and preset second calculation formula. Merge the target word frequency and the remaining word frequency to obtain multiple word frequencies, and merge the target inverse document frequency and the remaining inverse document frequency to obtain multiple inverse document frequencies; finally, the server screens out multiple target word corpora in the multiple word corpora whose word frequency is greater than or equal to the first set threshold and whose inverse document frequency is less than or equal to the second set threshold, and determines the text segments corresponding to the multiple target word corpora as processed data.

[0062] The server uses the above method to evaluate the importance of a word corpus for a document set or a single document in a corpus. The importance of a word corpus increases in direct proportion to the number of times it appears in a text segment, but at the same time decreases in inverse proportion to its frequency of occurrence in the corpus. That is to say, the more times a word corpus appears in a text segment and the fewer times it appears in all text segments, the more it can represent the center of the text segment. Here, the word frequency of the word corpus refers to the number of times a given word corpus appears in the corresponding text segment, and this number is usually normalized (generally, the word frequency is divided by the total number of words in the article) to prevent it from being biased towards long text segments. It should be noted that the same word corpus may have a higher word frequency in a long text segment than in a short text segment, regardless of whether the word corpus is a central word. The first pre-set calculation formula for calculating word frequency used here is:

[0063]

[0064] Among them, TF is the target word frequency of the candidate corpus, n is the number of times the candidate corpus appears in the text segment, s is the number of all word corpora in the text segment, and both n and s are positive integers.

[0065] After the server calculates the word frequency of the word corpus, it needs to calculate the inverse document frequency. Because if the fewer text segments contain a certain word corpus, the greater the inverse document frequency calculated by the pre-set second calculation formula, it indicates that the word corpus has category discrimination ability. The pre-set second calculation formula is:

[0066]

[0067] Among them, IDF is the target inverse document frequency of the candidate corpus, q is the number of text segments, z is the number of text segments where the candidate corpus exists, and both q and z are positive integers. It should be noted that in the pre-set second calculation formula, to prevent the denominator from being zero, the denominator in the formula is set to z + 1.

[0068] Finally, the server screens out multiple target word corpora with a word frequency greater than or equal to the first set threshold and an inverse document frequency less than or equal to the second set threshold from multiple word corpora, and determines the text segments corresponding to the multiple target word corpora as the processed data. In this way, the server screens and analyzes multiple text segments and obtains the processed data.

[0069] 205. Use the pre-set feature extractor to extract features from the processed data to obtain the feature vector of the processed data, and determine the feature label corresponding to the processed data according to the feature vector;

[0070] Specifically, the server first sends the processed data to a pre-set feature extractor, uses the pre-set feature extractor to extract features from the target word corpus in the processed data to obtain feature vectors; then the server calculates the similarity between the feature vectors and the label vectors to obtain a basic similarity; finally, the server selects the target similarity with the largest value of the basic similarity, and determines the pre-set label corresponding to the label vector for calculating the target similarity as the feature label corresponding to the processed data.

[0071] After obtaining the processed data, the server needs to extract representative feature labels from the processed data. First, the server uses the pre-set feature extractor to extract the feature vectors of the target word corpus in the processed data, and then calculates the similarity between the feature vectors and the label vectors to obtain a basic similarity. It should be noted that the number of label vectors here is multiple, so the calculated basic similarities are also multiple. In this application, the number of label vectors is not limited, and specifically, the number of label vectors can be set according to the actual situation.

[0072] Finally, the server will select the target similarity with the largest value of the basic similarity, and determine the pre-set label corresponding to the label vector for calculating the target similarity as the feature label corresponding to the processed data. Here, the label vector is the vector representation corresponding to the pre-set label, and each pre-set label has a unique corresponding label vector. The pre-set labels at least include platform characteristics, platform type, platform main business, platform competitiveness, platform customer group, platform language, address area, address location, address entity, address population, housing price corresponding to the address, convenience of the address, time seasonality, time specialness (holidays, shopping festivals), user repayment time, user salary payment time, commodity promotion time, commodity category, commodity name, commodity brand, commodity attributes, commodity customer group, commodity style, user demographic attributes, user family information, user economic level, user preference for purchasing commodities, user preference for using the platform, user consumption tendency, user loyalty, merchant platform, merchant location area, merchant customer group, merchant commodity richness, merchant market competitiveness, etc.

[0073] 206. Classify and analyze the feature labels through a pre-set clustering algorithm to obtain multi-dimensional weight labels. The weight labels at least include platform labels, address labels, time labels, commodity labels, user labels and merchant labels;

[0074] Specifically, the server first selects candidate tags from the feature tags using a preset clustering function; then, through a clustering algorithm, the server clusters the remaining tags with the candidate tags as the center to obtain grouped clustering tags, where the remaining tags are used to indicate the tags in the feature tags other than the candidate tags; finally, the server extracts the keywords of the grouped clustering tags and determines the keywords as the weight tags corresponding to the grouped clustering tags. The keyword is the central word of the grouped clustering tag, and the weight tags at least include platform tags, address tags, time tags, product tags, user tags, and merchant tags.

[0075] The server needs to perform statistical analysis on the feature tags. Therefore, the server uses the clustering function to cluster the feature tags. First, the server selects multiple candidate tags from the feature tags and clusters the remaining tags with the multiple candidate tags as the center to obtain grouped clustering tags, and determines the candidate tags as the weight tags. The clustering algorithm here is the k-means clustering algorithm, which is a commonly used technical means in the field and will not be elaborated here.

[0076] Furthermore, it should be noted that the weight tags here can be understood as the common features of the feature tags. For example, when the feature tags are platform characteristics, platform types, platform main businesses, platform competitiveness, platform customer groups, and platform languages, the corresponding weight tags are platform tags; when the feature tags are address regions, address locations, address entities, address population numbers, housing prices corresponding to the address, and the convenience of the address, the corresponding weight tags are address tags; when the feature tags are time seasonality, time specialness (holidays, shopping festivals), user repayment times, user salary payment times, and product promotion times, the corresponding weight tags are time tags; when the feature tags are product categories, product names, product brands, product attributes, product customer groups, and product styles, the corresponding weight tags are product tags; when the feature tags are user demographic attributes, user family information, user economic levels, user product purchase preferences, user platform usage preferences, user consumption tendencies, and user loyalty, the corresponding weight tags are user tags; when the feature tags are merchant platforms, merchant regions, merchant customer groups, merchant product richness, and merchant market competitiveness, the corresponding weight tags are merchant tags.

[0077] 207. Input the weight tags into a preset portrait generator, use the preset portrait generator to predict the weight tags to obtain predicted tags, and construct an express delivery industry portrait through the feature tags, weight tags, and predicted tags.

[0078] Specifically, the server first inputs the weight tags into a preset portrait generator, and uses the preset logistic regression model in the preset portrait generator to predict the weight tags to obtain the first predicted tag. Secondly, the server uses the preset product diffusion model in the preset portrait generator to predict the weight tags to obtain the second predicted tag. Then, the server uses the preset churn warning model in the preset portrait generator to predict the weight tags to obtain the third predicted tag. The server combines the first predicted tag, the second predicted tag, and the third predicted tag to obtain the predicted tag. Finally, the server inputs the feature tags, the weight tags, and the predicted tag into the system construction model in the preset portrait generator to generate an express delivery industry portrait.

[0079] After obtaining the weight tags, the server needs to predict the weight tags. First, the weight tags are input into a preset portrait generator, and the preset logistic regression model, the preset product diffusion model, and the preset churn warning model in the preset portrait generator are used to predict the weight tags in sequence. Through the above prediction steps, the corresponding first predicted tag, second predicted tag, and third predicted tag are obtained in sequence. The first predicted tag, the second predicted tag, and the third predicted tag are combined to obtain the predicted tag of the weight tags. The predicted tag here can be: the population attributes of users, the consumption ability of users, the default probability of users, the recent needs of users, the churn probability of users, etc.

[0080] Finally, the server generates an express delivery industry portrait by inputting the feature tags, the weight tags, and the predicted tag into the system construction model in the preset portrait generator. The express delivery industry portrait generated in this way makes the tags in the express delivery industry portrait more accurate and more in line with the actual situation of the express delivery industry.

[0081] In the embodiment of the present invention, by inputting the original express delivery data into a preset word segmenter and a preset feature extractor for processing, the feature tags corresponding to the original express delivery data are obtained, and then the feature tags are analyzed and predicted by a preset clustering algorithm and a preset portrait generator to generate a multi-dimensional express delivery industry portrait, improving the accuracy and efficiency of generating the express delivery industry portrait using the original express delivery data.

[0082] The generation method of the express delivery industry portrait in the embodiment of the present invention is described above. Next, the generation device of the express delivery industry portrait in the embodiment of the present invention will be described. Please refer to Figure 3 In one embodiment, the generation device of the express delivery industry portrait in the embodiment of the present invention includes:

[0083] A processing module 301, configured to obtain original express delivery data, segment the text segments in the original express delivery data based on a preset word segmenter to obtain word corpora, and obtain processed data by performing calculation processing on the word corpora;

[0084] A determination module 302, configured to extract features from the processed data by using the preset feature extractor to obtain a feature vector of the processed data, and determine a feature label corresponding to the processed data according to the feature vector;

[0085] A classification module 303, configured to perform classification analysis on the feature labels through a preset clustering algorithm to obtain multi-dimensional weight labels, where the weight labels at least include a platform label, an address label, a time label, a commodity label, a user label, and a merchant label;

[0086] A generation module 304, configured to input the weight labels into a preset portrait generator, use the preset portrait generator to predict the weight labels to obtain predicted labels, and construct an express industry portrait through the feature labels, the weight labels, and the predicted labels.

[0087] In an embodiment of the present invention, by inputting the original express data into a preset tokenizer and a preset feature extractor for processing, a feature label corresponding to the original express data is obtained, and then through a preset clustering algorithm and a preset portrait generator, the feature labels are analyzed and predicted to generate a multi-dimensional express industry portrait, improving the accuracy and efficiency of generating an express industry portrait using the original express data.

[0088] Please refer to Figure 4 , another embodiment of the express industry portrait generation device in the embodiment of the present invention includes:

[0089] A processing module 301, configured to obtain original express data, segment the paragraphs in the original express data based on a preset tokenizer to obtain a word corpus, and obtain processed data by performing calculation processing on the word corpus;

[0090] A determination module 302, configured to extract features from the processed data by using the preset feature extractor to obtain a feature vector of the processed data, and determine a feature label corresponding to the processed data according to the feature vector;

[0091] A classification module 303, configured to perform classification analysis on the feature labels through a preset clustering algorithm to obtain multi-dimensional weight labels, where the weight labels at least include a platform label, an address label, a time label, a commodity label, a user label, and a merchant label;

[0092] A generation module 304, configured to input the weight labels into a preset portrait generator, use the preset portrait generator to predict the weight labels to obtain predicted labels, and construct an express industry portrait through the feature labels, the weight labels, and the predicted labels.

[0093] Optionally, the processing module 301 includes:

[0094] An acquisition unit 3011, configured to acquire original express delivery data and transmit the original express delivery data to a preset word segmenter;

[0095] A segmentation unit 3012, configured to segment the language segments in the original express delivery data into multiple word corpora in the preset word segmenter and count the number of the multiple word corpora, where the word corpora are words or phrases existing in a standard dictionary;

[0096] A statistics unit 3013, configured to use a preset statistical function to count the frequencies of occurrence of each word corpus in the original express delivery data to obtain multiple basic frequencies;

[0097] A determination unit 3014, configured to calculate the number of occurrences of the corresponding word corpus in the language segment through each basic frequency to obtain multiple word frequencies, calculate the inverse document frequency of each word corpus to obtain multiple inverse document frequencies, and determine multiple target word corpora according to the multiple word frequencies and the multiple inverse document frequencies to obtain processed data.

[0098] Optionally, the determination unit 3014 is specifically configured to:

[0099] Obtain candidate corpora in the word corpus, calculate the number of occurrences of the candidate corpus in the language segment through the basic frequency corresponding to the candidate corpus and a preset first calculation formula to obtain a target word frequency, where the preset first calculation formula is:

[0100]

[0101] where TF is the target word frequency of the candidate corpus, n is the number of occurrences of the candidate corpus in the language segment, s is the number of all word corpora in the language segment, and both n and s are positive integers;

[0102] Calculate the inverse document frequency of the candidate corpus by using a preset second calculation formula, where the preset second calculation formula is:

[0103]

[0104] where IDF is the target inverse document frequency of the candidate corpus, q is the number of language segments, z is the number of language segments where the candidate corpus exists, and both q and z are positive integers;

[0105] Obtain the remaining corpora in the word corpus except the candidate corpus, calculate the remaining word frequency and the remaining inverse document frequency of the remaining corpus through the preset first calculation formula and the preset second calculation formula, merge the target word frequency and the remaining word frequency to obtain multiple word frequencies, and merge the target inverse document frequency and the remaining inverse document frequency to obtain multiple inverse document frequencies;

[0106] Filter out multiple target word corpora with a word frequency greater than or equal to the first set threshold and a reverse corpus frequency less than or equal to the second set threshold from multiple word corpora, and determine the paragraphs corresponding to the multiple target word corpora as processed data.

[0107] Optionally, the determination module 302 is specifically configured to:

[0108] Send the processed data to a pre-set feature extractor, and use the pre-set feature extractor to extract features from the target word corpora in the processed data to obtain feature vectors;

[0109] Calculate the similarity between the feature vectors and the label vectors to obtain a basic similarity;

[0110] Select the target similarity with the largest value of the basic similarity, and determine the pre-set label corresponding to the label vector for calculating the target similarity as the feature label corresponding to the processed data.

[0111] Optionally, the classification module 303 is specifically configured to:

[0112] Select candidate labels from the feature labels by using a pre-set clustering function;

[0113] Through a clustering algorithm, cluster the remaining labels with the candidate labels as the center to obtain grouped clustering labels, where the remaining labels are used to indicate the labels in the feature labels other than the candidate labels;

[0114] Extract the keywords of the grouped clustering labels, and determine the keywords as the weight labels corresponding to the grouped clustering labels. The keyword is the central word of the grouped clustering label, and the weight labels at least include platform labels, address labels, time labels, product labels, user labels, and merchant labels.

[0115] Optionally, the generation module 304 is specifically configured to:

[0116] Use the pre-set product diffusion model in the pre-set portrait generator to predict the weight labels to obtain second prediction labels;

[0117] Use the pre-set churn warning model in the pre-set portrait generator to predict the weight labels to obtain third prediction labels;

[0118] Merge the first prediction label, the second prediction label, and the third prediction label to obtain a prediction label;

[0119] Input the feature label, the weight label, and the prediction label into the system construction model in the pre-set portrait generator to generate an express delivery industry portrait.

[0120] In the embodiments of the present invention, by inputting the original express delivery data into a preset word segmenter and a preset feature extractor for processing, a feature label corresponding to the original express delivery data is obtained. Then, through a preset clustering algorithm and a preset portrait generator, the feature label is analyzed and predicted to generate a multi-dimensional portrait of the express delivery industry, improving the accuracy and efficiency of generating the portrait of the express delivery industry using the original express delivery data.

[0121] Above Figure 3 And Figure 4 The generation device of the portrait of the express delivery industry in the embodiments of the present invention is described in detail from the perspective of modular functional entities. Next, the generation device of the portrait of the express delivery industry in the embodiments of the present invention is described in detail from the perspective of hardware processing.

[0122] Figure 5 FIG. is a schematic structural diagram of a generation device of a portrait of the express delivery industry provided by an embodiment of the present invention. The generation device 500 of the portrait of the express delivery industry may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPU) 510 (for example, one or more processors) and a memory 520, and one or more storage media 530 (for example, one or more mass storage devices) for storing application programs 533 or data 532. Among them, the memory 520 and the storage media 530 may be transient storage or persistent storage. The program stored in the storage media 530 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the generation device 500 of the portrait of the express delivery industry. Further, the processor 510 may be configured to communicate with the storage media 530 and execute a series of instruction operations in the storage media 530 on the generation device 500 of the portrait of the express delivery industry.

[0123] The generation device 500 of the portrait of the express delivery industry may further include one or more power supplies 540, one or more wired or wireless network interfaces 550, one or more input / output interfaces 560, and / or one or more operating systems 531, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand that Figure 5 The shown structural diagram of the generation device of the portrait of the express delivery industry does not limit the generation device of the portrait of the express delivery industry, and may include more or fewer components than shown, or combine some components, or have different component arrangements.

[0124] The present invention also provides a device for generating a portrait of the express delivery industry. The computer device includes a memory and a processor. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by the processor, the processor executes the steps of the method for generating a portrait of the express delivery industry in the above-mentioned various embodiments.

[0125] The present invention also provides a computer-readable storage medium. The computer-readable storage medium can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the method for generating a portrait of the express delivery industry.

[0126] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0127] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0128] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for generating a portrait of the express delivery industry, characterized in that, The method for generating the express delivery industry portrait includes: Obtain the original express delivery data, segment the text segments in the original express delivery data based on a preset word segmenter to obtain word corpora, and obtain processed data by performing calculation processing on the word corpora. Use the preset feature extractor to extract features from the processed data to obtain the feature vectors of the processed data, and determine the feature labels corresponding to the processed data according to the feature vectors. Perform classification analysis on the feature labels through a preset clustering algorithm to obtain multi-dimensional weight labels, and the weight labels at least include platform labels, address labels, time labels, commodity labels, user labels, and merchant labels. Input the weight labels into a preset portrait generator, use the preset portrait generator to predict the weight labels to obtain prediction labels, and construct an express delivery industry portrait through the feature labels, the weight labels, and the prediction labels. The step of obtaining the original express delivery data, segmenting the text segments in the original express delivery data based on a preset word segmenter to obtain word corpora, and obtaining processed data by performing calculation processing on the word corpora includes: Obtain the original express delivery data and transmit the original express delivery data to a preset word segmenter. Segment the text segments in the original express delivery data into multiple word corpora in the preset word segmenter, and count the number of the multiple word corpora. The word corpora are words or phrases existing in a standard dictionary. Use a preset statistical function to count the frequencies of each word corpus in the original express delivery data to obtain multiple basic frequencies. Calculate the number of times each corresponding word corpus appears in the text segment through each basic frequency to obtain multiple word frequencies, calculate the inverse document frequency of each word corpus to obtain multiple inverse document frequencies, and determine multiple target word corpora according to the multiple word frequencies and the multiple inverse document frequencies to obtain processed data. The step of calculating the number of times each corresponding word corpus appears in the text segment through each basic frequency to obtain multiple word frequencies, calculating the inverse document frequency of each word corpus to obtain multiple inverse document frequencies, and determining multiple target word corpora according to the multiple word frequencies and the multiple inverse document frequencies to obtain processed data includes: Obtain the candidate corpora in the word corpora, calculate the number of times the candidate corpora appear in the text segment through the basic frequency corresponding to the candidate corpora and a preset first calculation formula to obtain the target word frequency. The preset first calculation formula is: where TF is the target word frequency of the candidate corpus, n is the number of times the candidate corpus appears in the text segment, s is the number of all word corpora in the text segment, and both n and s are positive integers. Use a preset second calculation formula to calculate the inverse document frequency of the candidate corpus. The preset second calculation formula is: where IDF is the target inverse document frequency of the candidate corpus, q is the number of text segments, z is the number of text segments where the candidate corpus exists, and both q and z are positive integers. Obtain the remaining corpus in the word corpus except the candidate corpus, calculate the remaining word frequency and the remaining reverse corpus frequency of the remaining corpus through the preset first calculation formula and the preset second calculation formula, merge the target word frequency and the remaining word frequency to obtain multiple word frequencies, and merge the target reverse corpus frequency and the remaining reverse corpus frequency to obtain multiple reverse corpus frequencies; Screen out multiple target word corpora in multiple word corpora whose word frequency is greater than or equal to the first set threshold and whose reverse corpus frequency is less than or equal to the second set threshold, and determine the text segments corresponding to the multiple target word corpora as the processed data.

2. The method for generating a portrait of the express delivery industry according to claim 1, characterized in that, The feature extraction of the processed data by using the preset feature extractor to obtain the feature vector of the processed data, and determining the feature label corresponding to the processed data according to the feature vector includes: Send the processed data to the preset feature extractor, and use the preset feature extractor to perform feature extraction on the target word corpus in the processed data to obtain a feature vector; Calculate the similarity between the feature vector and the label vector to obtain the basic similarity; Select the target similarity with the largest value of the basic similarity, and determine the preset label corresponding to the label vector for calculating the target similarity as the feature label corresponding to the processed data.

3. The method for generating a portrait of the express delivery industry according to claim 1, characterized in that, The classification analysis of the feature labels by using the preset clustering algorithm to obtain multi-dimensional weight labels, and the weight labels at least include platform labels, address labels, time labels, commodity labels, user labels and merchant labels, including: Use the preset clustering function to select candidate labels from the feature labels; Through the clustering algorithm, cluster the remaining labels with the candidate labels as the center to obtain grouped clustering labels, where the remaining labels are used to indicate the labels in the feature labels other than the candidate labels; Extract the keywords of the grouped clustering labels, and determine the keywords as the weight labels corresponding to the grouped clustering labels. The keywords are the central words of the grouped clustering labels, and the weight labels at least include platform labels, address labels, time labels, commodity labels, user labels and merchant labels.

4. The method for generating a portrait of the express delivery industry according to any one of claims 1-3, characterized in that, Input the weight labels into the preset portrait generator, use the preset portrait generator to predict the weight labels to obtain predicted labels, and construct an express delivery industry portrait through the feature labels, the weight labels and the predicted labels, including: Input the weight labels into the preset portrait generator, and use the preset logistic regression model in the preset portrait generator to predict the weight labels to obtain the first predicted label; Use the preset product diffusion model in the preset portrait generator to predict the weight labels to obtain the second predicted label; Use the preset churn warning model in the preset portrait generator to predict the weight labels to obtain the third predicted label; Merge the first predicted label, the second predicted label and the third predicted label to obtain the predicted label; Input the feature label, the weight label, and the prediction label into the system construction model in the preset portrait generator to generate an express delivery industry portrait.

5. A device for generating a portrait of the express delivery industry, characterized in that,The generating device for the express delivery industry portrait includes: A processing module, configured to obtain original express delivery data, segment the text segments in the original express delivery data based on a preset word segmenter to obtain word corpora, and obtain processed data by performing calculation processing on the word corpora. A determining module, configured to use the preset feature extractor to extract features from the processed data to obtain a feature vector of the processed data, and determine a feature label corresponding to the processed data according to the feature vector. A classification module, configured to perform classification analysis on the feature labels through a preset clustering algorithm to obtain multi-dimensional weight labels, where the weight labels at least include a platform label, an address label, a time label, a commodity label, a user label, and a merchant label. A generating module, configured to input the weight labels into a preset portrait generator, use the preset portrait generator to predict the weight labels to obtain prediction labels, and construct an express delivery industry portrait through the feature labels, the weight labels, and the prediction labels. The processing module includes: An obtaining unit, configured to obtain original express delivery data and transmit the original express delivery data to a preset word segmenter. A segmenting unit, configured to segment the text segments in the original express delivery data into multiple word corpora in the preset word segmenter and count the number of the multiple word corpora, where the word corpora are words or phrases existing in a standard dictionary. A statistical unit, configured to use a preset statistical function to count the frequency of each word corpus appearing in the original express delivery data to obtain multiple basic frequencies. A determining unit, configured to calculate the number of times each word corpus appears in the text segment through each basic frequency to obtain multiple word frequencies, calculate the inverse document frequency of each word corpus to obtain multiple inverse document frequencies, and determine multiple target word corpora according to the multiple word frequencies and the multiple inverse document frequencies to obtain processed data. The determining unit is further configured to obtain candidate corpora in the word corpora, calculate the number of times the candidate corpora appear in the text segment through the basic frequency corresponding to the candidate corpora and a preset first calculation formula to obtain target word frequencies, and the preset first calculation formula is: where TF is the target word frequency of the candidate corpus, n is the number of times the candidate corpus appears in the text segment, s is the number of all word corpora in the text segment, and n and s are both positive integers. Calculate the inverse document frequency of the candidate corpus by using a preset second calculation formula, and the preset second calculation formula is: where IDF is the target inverse document frequency of the candidate corpus, q is the number of text segments, z is the number of text segments where the candidate corpus exists, and q and z are both positive integers. Obtain the remaining corpus in the word corpus except for the candidate corpus, calculate the remaining word frequency and the remaining reverse corpus frequency of the remaining corpus through the preset first calculation formula and the preset second calculation formula, merge the target word frequency and the remaining word frequency to obtain multiple word frequencies, and merge the target reverse corpus frequency and the remaining reverse corpus frequency to obtain multiple reverse corpus frequencies; Screen out multiple target word corpora in multiple word corpora whose word frequency is greater than or equal to the first set threshold and whose reverse corpus frequency is less than or equal to the second set threshold, and determine the paragraphs corresponding to the multiple target word corpora as the processed data.

6. An apparatus for generating a portrait of the express delivery industry, characterized in that, The generating device for the express delivery industry portrait includes: a memory and at least one processor, and instructions are stored in the memory; The at least one processor calls the instructions in the memory so that the generating device for the express delivery industry portrait executes the method for generating the express delivery industry portrait according to any one of claims 1-4.

7. A computer-readable storage medium having instructions stored thereon, characterized in that, When the instructions are executed by the processor, the method for generating the express delivery industry portrait according to any one of claims 1-4 is implemented.

Citation Information

Patent Citations

  • User portrait construction system

    CN107578292A

  • Tourist ticket product portrait generation method

    CN110910175A