Business opportunity mining method and device based on big data and medium
By leveraging big data and deep learning technologies, multi-dimensional enterprise profiles are constructed, solving the data limitations and real-time issues of traditional business opportunity mining methods. This enables efficient and automated business opportunity mining, improving the accuracy of business opportunity identification and the timeliness of decision-making.
Patent Information
- Application Number
- CN202511299815.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2026-02-13
AI Technical Summary
Traditional business opportunity discovery methods rely on limited data sources, have single analytical dimensions, lack real-time capabilities, and involve a lot of manual intervention, resulting in low efficiency and poor accuracy in business opportunity discovery, making it difficult to cope with the rapidly changing market environment.
By employing big data analytics and deep learning technologies, and through data preprocessing and feature extraction, multi-dimensional enterprise profiles are constructed using enterprise user profile models and product representation models. These profiles are then combined with deep neural network models to uncover business opportunities, enabling automated and real-time analysis.
It improves the accuracy and efficiency of business opportunity discovery, can capture market dynamics in real time, provide efficient business opportunity suggestions, reduce labor and time costs, and enhance the timeliness of decision-making.
Smart Images

Figure CN121526653A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of deep learning and big data technology, and in particular to a method, device and medium for mining business opportunities based on big data. Background Technology
[0002] With the rapid development of the internet and information technology, the global volume of data has grown exponentially, ushering in the era of big data. Various enterprises, government agencies, and research organizations have accumulated vast amounts of data resources, covering multiple dimensions such as consumer behavior, market dynamics, competitor activities, and public opinion. Extracting valuable information from this massive amount of data, especially identifying potential business opportunities, has become a key factor for enterprises to gain a competitive advantage.
[0003] Traditional methods of identifying business opportunities rely on expert experience, market research reports, and simple statistical analysis of historical data. These methods have the following limitations: Limited data sources: Traditional methods often rely on limited data sources, such as internal sales data or a small number of market research reports, which are difficult to fully reflect the dynamic changes in the market.
[0004] Limited analytical dimensions: Traditional methods focus primarily on quantitative analysis and cannot effectively combine multi-dimensional data (such as social media, consumer feedback, macroeconomic indicators, etc.) for comprehensive analysis, making it difficult to identify business opportunities hidden in complex data relationships.
[0005] Insufficient real-time capability: Traditional market analysis and business opportunity discovery are mostly conducted periodically, which cannot cope with the rapidly changing market environment and often misses the opportunity to capture business opportunities in a timely manner.
[0006] High degree of human intervention and low efficiency: Traditional business opportunity discovery methods require a lot of human intervention, the process is complex and time-consuming, and it is easily affected by human factors, which limits the efficiency and accuracy of business opportunity discovery.
[0007] Against this backdrop, business opportunity mining methods based on big data technology have gradually become a research hotspot. By introducing technologies such as big data analytics, machine learning, and natural language processing, potential business opportunities can be automatically mined from massive amounts of heterogeneous data. This method can not only integrate data from multiple sources, improving the comprehensiveness and accuracy of business opportunity identification, but also analyze data dynamic changes in real time, enabling rapid response to market demands.
[0008] However, existing big data opportunity mining methods and equipment still face many challenges in application, such as low data processing efficiency, insufficient generalization ability of analytical models, difficulty in meeting personalized user needs, and data security and privacy issues. Therefore, there is an urgent need to propose a new big data-based opportunity mining method, equipment, and storage medium to improve the accuracy, efficiency, and user experience of opportunity mining, helping enterprises gain a favorable position in a highly competitive market environment. Summary of the Invention
[0009] To address the aforementioned problems, this invention proposes a method, equipment, and medium for business opportunity mining based on big data.
[0010] The first aspect of this application provides a business opportunity mining method based on big data, comprising the following steps: Obtain the transactional dataset A and the enterprise basic dataset B; The datasets A and B are preprocessed to obtain preprocessed data. The pre-processed data was used to pre-train the enterprise user profiling model ECPM and the product representation model PRM. We construct multi-dimensional enterprise profiles using ECPM and PRM, and train a deep neural network model, ResDeepFM, using enterprise basic data; Business opportunity mining is performed using a deep neural network model to obtain business opportunity mining results.
[0011] A second aspect of this application provides a terminal including a processor, an input device, an output device, and a memory, wherein the processor, input device, output device, and memory are interconnected, wherein the memory is used to store a computer program, the computer program including program instructions, and the processor is configured to invoke the program instructions to execute the step instructions as described in the first aspect of this application.
[0012] A third aspect of this application provides a computer-readable storage medium storing a computer program for electronic data interchange, wherein the computer program causes a computer to perform some or all of the steps described in the first aspect of this application.
[0013] A fourth aspect of this application provides a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps described in the first aspect of this application. The computer program product may be a software installation package.
[0014] The beneficial effects of this invention are: 1. This invention uses data cleaning, preprocessing, and feature extraction techniques to ensure the high quality and consistency of input data, thereby improving the accuracy and reliability of business opportunity discovery; 2. This invention utilizes advanced natural language processing and deep learning technologies to efficiently represent (embedding) product and enterprise profiles. This representation method can capture the complex relationships and potential interaction patterns between items, helping to uncover potential related business opportunities.
[0015] 3. This invention utilizes deep learning models, especially the multi-layered structure of neural networks. This automated feature learning reduces the complexity and uncertainty of human intervention, ensuring that the model can extract the most relevant features from the data, improving the accuracy and effectiveness of business opportunity mining, while reducing the human and time costs in the model development process. 4. This invention employs advanced big data analysis algorithms and deep learning models, enabling it to automatically identify potential business opportunities from massive amounts of data. Through comprehensive analysis of multi-dimensional data such as market trends, user behavior, and competitive landscape, it accurately identifies growth points and market gaps, providing enterprises with actionable business opportunity suggestions. 5. This invention supports real-time data updates and analysis, enabling timely capture of dynamic market changes and helping businesses respond quickly to market fluctuations. This real-time advantage ensures that businesses do not miss any potential opportunities, significantly improving the timeliness of decision-making. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below.
[0017] Figure 1 This is an overall flowchart of a business opportunity mining method based on big data according to the present invention; Figure 2 This is the specific training process of the Enterprise User Profile Model (ECPM) and the Product Representation Model (PRM) in this invention; Figure 3 This describes the specific training process of ResDeepFM in this invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] This invention proposes a business opportunity mining method, equipment, and medium based on big data, the process of which is as follows: Figure 1 As shown. Includes: 101. Collect information and transaction data from enterprise users to generate transaction dataset A and enterprise basic dataset B. The steps include: S1. Extract all enterprise transaction form data and enterprise user information tables from the business database.
[0020] S2. Extract enterprise user information, including but not limited to basic information such as enterprise ID, contact information, registration time, and industry classification, to obtain the enterprise basic dataset B. Extract data from the enterprise user transaction data table, including but not limited to key information such as enterprise ID, information on traded items (such as item name, quantity, price, etc.), and transaction time, to obtain the transaction dataset A.
[0021] S3. Examine datasets A and B row by row for missing and outlier values. If missing values are found, fill them with the mean of the feature; if outliers are found, delete the relevant rows to ensure data accuracy and consistency; convert date data to timestamp format and extract key date information such as year, month, and day for subsequent analysis and modeling; label and encode categorical features.
[0022] 102. Perform data preprocessing on datasets A and B to obtain preprocessed data, and store it in the data center. The steps include: A1. Concatenate datasets A and B according to enterprise ID to form dataset C, which consists of "seller-seller information-buyer-buyer information-traded item-transaction information", and sort it according to transaction time.
[0023] A2. Aggregate C by enterprise ID to obtain "Enterprise-Transaction Item Sequence Data" I, and aggregate C by transaction item to obtain "Transaction Item-Enterprise Sequence Data" H.
[0024] A3. Store A, B, C, H, and I in a MySQL data center. MySQL is a computer program that can run on a computer-readable storage medium, and data can be stored on such a medium through the program.
[0025] 103. Using the preprocessed data, pre-train the Enterprise User Profile Model (ECPM) and the Product Representation Model (PRM). Taking the training of the Enterprise User Profile Model (ECPM) as an example, the training process is as follows: Figure 2 Similarly, the Enterprise User Profile Model (ECPM) can be trained, and the steps include: B1. Read dataset H, which contains P data entries. Taking one data entry O as an example, O has n elements.
[0026] B2. Utilize the Transformer-Encoder architecture and employ a non-replacement MLM mask pre-training method to obtain the Enterprise User Profile Model (ECPM). The steps include: B3. Input Data ,in This represents the i-th element of the input sequence o, with a randomly selected position. The location of the mask.
[0027] Where p is the probability of each position being selected as a mask position, and its value is 0.3. It follows a Bernoulli distribution, and the probability of the result being 1 is... The probability of the result being 0 is .
[0028] For each i in the selected position S, perform a masking operation to... Replace with mask marker The formula is as follows: B4. Input the Transformer-Encoder model, which predicts the masked elements based on the context elements. Specifically, for each masked position... The model generates a probability distribution. The probability distribution represents the probability of all elements appearing at that position.
[0029] B5. Calculate the loss and backpropagate. When calculating the loss, due to the time-sensitive nature of opportunity discovery, the loss function uses the most recent weighted cross-entropy. The calculation process is as follows: S51. Take out dataset C, and calculate the occurrence count of each company to obtain the "Company ID - Occurrence Count" data T; take the data from the most recent 3 months of dataset C to obtain... ,exist The above data, "Company ID - Number of Occurrences", is obtained by statistically analyzing the occurrence counts of each company. S52. Calculate the recent weight (Recent_weight): Where id is one of all enterprise IDs, Σ represents iterating through all ids, calculating the weight for each id, and finally concatenating all the results to obtain Recent_weight.
[0030] S53. Calculate the loss using the following formula: in This represents the i-th element of the input sequence O. express After one-hot encoding, the value of the id-th element is either 0 or 1. This represents the id-th element of the predicted distribution. Let d represent the weight of the id-th element, and Σ be the summation symbol.
[0031] By assigning higher weights to recent data, recent-weighted cross-entropy effectively reduces the interference of outdated information on model predictions, allowing the model to pay more attention to and value the latest market dynamics and user behavior, thereby better representing products and enterprises and improving the accuracy of current business opportunities.
[0032] 104. A multi-dimensional enterprise profile is constructed using ECPM and PRM, and a deep neural network model, ResDeepFM, is trained using the enterprise's basic data to obtain the target model. The training process is as follows: Figure 3 As shown, the steps include: C1. Input the enterprise ID and the transaction item ID into the models EPCM and PRM respectively to obtain the representation information of each enterprise and each item. Concatenate these representations to obtain the Embedding-feature. Extract basic information from dataset C to form dataset X.
[0033] C2. Concatenate the Embedding_feature with X to obtain E, and input it into the Residual connections-DNN. The calculation formula is as follows: Where linear represents a linear layer. It is the output of Residual connections-DNN.
[0034] C3. Input X into the FM section, and the calculation formula is as follows: Where l is the number of features in dataset X, and j, k are used to label that this is the j-th, k-th feature of X. It is a global bias term. Is it related to the j-th feature? Associated weights It is a feature The latent vectors are used to capture the interactions between features. Representation of features The inner product of the latent vectors. y is the final prediction result of the model. These are two learnable parameters, initially set to 0.5, 0.5. C4. Perform backpropagation and update the parameters using gradient descent until the model converges.
[0035] Residual connections allow information to be passed more directly within the network, thus accelerating model training convergence. By reducing the loss of information between layers, the network can learn effective feature representations more quickly, thereby improving training efficiency. Simultaneously, residual connections allow the network to learn more complex feature combinations because they not only capture layer-by-layer feature changes but also retain the original information of the input. This combination and fusion of information enhances the network's expressive power.
[0036] 105. Opportunity mining is performed using the target model to obtain the results.
[0037] The data to be detected is input into the target model for prediction processing to obtain the business opportunity mining results. One specific implementation also includes a business opportunity mining system. This system comprises a data processing module that preprocesses the collected data and stores it in a data center; a business opportunity mining module that trains models and makes predictions on a test set by processing the data; a recommendation module that displays and pushes feasible business opportunities based on the prediction results; a user feedback module that allows users to evaluate and provide feedback on the pushed business opportunities; and a security and privacy protection module that equips the device with data encryption, access control, and log auditing functions to ensure the security of user data and business information.
[0038] The above embodiments should be understood as illustrative only and not as limiting the scope of protection of the present invention. After reading the description of the present invention, those skilled in the art can make various alterations or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.
Claims
1. A big data-based business opportunity mining method, characterized by, The method comprises the following steps: Obtaining a transaction dataset A and a business basic dataset B; Data preprocessing is performed on the data sets A and B to obtain preprocessed data; The preprocessed data is used to pretrain a business user portrait model ECPM and a commodity representation model PRM; A multi-dimensional business portrait is constructed by using the ECPM and the PRM, and a deep neural network model ResDeepFM is trained in combination with business basic data to obtain a target model; Opportunity mining is performed by using the target model to obtain an opportunity mining result. 2.The big data-based business opportunity mining method of claim 1, wherein, Obtaining a transaction dataset A and a business basic dataset B comprises: Extracting all enterprise transaction form data and enterprise user information tables from a business database; Extracting enterprise user information from the enterprise user information table to obtain the business basic dataset B, wherein the enterprise user information includes enterprise ID, contact information, registration time, industry classification and other basic information to obtain the business basic dataset B; Extracting a transaction dataset A from the enterprise user transaction data table, wherein the transaction dataset A includes enterprise ID, transaction item information and transaction time. 3.The big data-based business opportunity mining method of claim 1, wherein, The data preprocessing of the data sets A and B comprises: Splicing the data sets A and B according to the enterprise ID to form a data set C of "seller-seller information-buyer-buyer information-transaction item-transaction information", and sorting the data set C according to the transaction time; Aggregating the C according to the enterprise ID to obtain "enterprise-transaction item sequence data" I, and aggregating the C according to the transaction item to obtain "transaction item-enterprise sequence data" H; Storing the A, B, C, H and I in a MySql data center. 4.The big data-based business opportunity mining method of claim 1, wherein, The preprocessed data is used to pretrain a business user portrait model ECPM, which comprises: Reading the data set H, a total of P data, taking one data O as an example, O has a total of n elements; Using the Transformer-Encoder structure and the non-replacement MLM mask pretraining method to obtain the business user portrait model ECPM; Using the Transformer-Encoder structure and the non-replacement MLM mask pretraining method to obtain the business user portrait model ECPM comprises: Input data , wherein represents the i-th element of the input sequence O, a set of positions are randomly selected as the positions of the mask; ; where p is the probability of each position being selected as a mask position, size is 0.3, is a Bernoulli distribution with probability of 1 and probability of 0 ; Masking for each i in the selected positions S will be replaced by the mask marker , which is given by the formula The input Transformer-Encoder model predicts the masked elements by the context elements; wherein, for each masked position , the model generates a probability distribution , the probability distribution represents the probability of all elements appearing at this position; Calculating the loss and backpropagating, wherein the loss function adopts the nearest weighted cross entropy when calculating the loss. 5.The big data-based business opportunity mining method of claim 4, wherein, The nearest weighted cross entropy calculation comprises: Take dataset C and count the occurrences of each company to obtain the "Company ID - Occurrence Count" data T; take the data from the most recent 3 months of dataset C to obtain... ,exist The above data, "Company ID - Number of Occurrences", is obtained by statistically analyzing the occurrence counts of each company. ; Calculating the nearest weight Recent_weight: Where id is one of all enterprise IDs, Σ represents traversal, the weight of each id is calculated, and finally all results are spliced to obtain Recent_weight; Calculating loss, the formula is as follows: wherein denotes the i-th element of the input sequence O, denotes denotes the i-th element after one-hot, value 0 or 1, denotes the i-th element of the predicted distribution, denotes the i-th weight, and ∑ is the summation sign over all i. 6.The big data-based business opportunity mining method of claim 4, wherein, A multi-dimensional business portrait is constructed by using the ECPM and the PRM, and a deep neural network model ResDeepFM is trained in combination with business basic data to obtain a target model, which comprises: Inputting the enterprise ID and the transaction item ID into the model EPCM and the PRM respectively to obtain the representation information of each enterprise and each item, splicing the representation information to obtain Embedding-feature, and extracting the basic information from the data set C to form a data set X; The Embedding_feature is spliced with X to obtain E, which is input into a Residual connections-DNN, and the calculation formula is as follows: ; ; ; wherein linear represents a linear layer, is the output of the Residual connections-DNN; X is input into the FM part, and the calculation formula is as follows: Where l is the number of features in dataset X, and j, k are used to label that this is the j-th, k-th feature of X. It is a global bias term. Is it related to the j-th feature? Associated weights It is a feature The latent vectors are used to capture the interactions between features. Representation of features The inner product of the latent vectors, where y is the final prediction result of the model. These are two learnable parameters, with initial values of 0.5 and 0.
5. Back propagation is performed, and the gradient descent method is used for parameter updating until the model converges.
7. An apparatus, comprising: The computer readable storage medium stores a computer program, and the computer program includes program instructions.
8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program includes program instructions.