Big data-based technological enterprise portrait information pushing system
By combining big data-based cloud computing and edge computing modules with sharding encryption technology and an improved butterfly optimization algorithm, the data storage and push scheme of the technology enterprise profile information push system has been optimized, solving the efficiency and timeliness problems of the existing system and realizing efficient, secure, and personalized information push.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- KECHUANGTONG CHENGDU CO LTD
- Filing Date
- 2026-03-19
- Publication Date
- 2026-05-29
Smart Images

Figure CN121887857B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information push, and in particular to a big data-based information push system for profiling technology companies. Background Technology
[0002] With the rapid development of technology and the wave of digital transformation, users' demand for accurate and timely information is increasing daily to assist decision-making, grasp market dynamics, and promote innovative development. However, existing information push systems for technology companies are insufficient to meet actual needs in terms of efficiency and timeliness.
[0003] Traditional systems often rely on centralized data processing architectures, where all data is aggregated and processed by a single central server. However, with the exponential growth of data volume in technology companies, encompassing massive amounts of information across multiple dimensions such as business operations, market dynamics, and technological research and development, centralized architectures face immense processing pressure, are prone to latency, and suffer from low push efficiency, failing to quickly respond to enterprise needs.
[0004] Therefore, there is a need to provide a big data-based information push system for technology companies to improve the efficiency and timeliness of push notifications. Summary of the Invention
[0005] This invention provides a big data-based technology enterprise profile information push system, comprising: a cloud computing module for generating multiple technology enterprise profiles; a profile storage module for storing multiple technology enterprise profiles; and an edge computing module including multiple edge nodes. The cloud computing module is further configured to acquire push impact information from at least one user terminal, determine the target technology enterprise profile and optimal push scheme for the user terminal based on the push impact information from the user terminal, and push the target technology enterprise profile to the user terminal through the edge computing module.
[0006] Furthermore, the portrait storage module is also used for: for each technology-based enterprise portrait, segmenting the technology-based enterprise portrait to generate multiple portrait fragments and segment identifiers for the technology-based enterprise portrait, wherein the segment identifier is generated based on the hash value and segment position of the portrait fragment; for each technology-based enterprise portrait, encrypting multiple portrait fragments of the technology-based enterprise portrait based on the hash value corresponding to the segment identifier of the technology-based enterprise portrait to generate multiple encrypted portrait fragments; encrypting the segment identifier of the technology-based enterprise portrait using the key corresponding to the technology-based enterprise portrait, and storing the encrypted multiple portrait fragments and the encrypted segment identifier.
[0007] Furthermore, the profile storage module is also used to: search for similar technology enterprise profiles from the stored technology enterprise profiles; obtain historical query information of similar technology enterprise profiles; determine the optimal distributed storage scheme based on the historical query information of similar technology enterprise profiles; and store multiple encrypted profile fragments and encrypted fragment identifiers based on the optimal distributed storage scheme.
[0008] Furthermore, the profile storage module is also used for: predicting the query frequency of technology enterprise profiles based on historical query information of similar technology enterprise profiles; determining the optimal number of storage nodes based on the query frequency of technology enterprise profiles; generating multiple distributed storage schemes based on the optimal number of storage nodes, wherein each distributed storage scheme includes multiple storage nodes for storing encrypted profile fragments and encrypted fragment identifiers; for each distributed storage scheme, determining the evaluation score corresponding to the distributed storage scheme based on the historical query information of technology enterprise profiles already stored in each storage node included in the distributed storage scheme; and determining the optimal distributed storage scheme based on the evaluation score corresponding to each distributed storage scheme.
[0009] Furthermore, the profile storage module is also used to: determine the average access frequency of each storage node in the distributed storage scheme based on the historical query information of the profiles of technology-based enterprises stored in each storage node of the distributed storage scheme; determine the access correlation coefficient between any two storage nodes in the distributed storage scheme based on the historical query information of the profiles of technology-based enterprises stored in each storage node of the distributed storage scheme; and calculate the evaluation score corresponding to the distributed storage scheme based on the average access frequency of each storage node in the distributed storage scheme and the access correlation coefficient between any two storage nodes in the distributed storage scheme.
[0010] Furthermore, the cloud computing module is also used to: acquire historical push information from multiple user terminals, wherein the historical push information includes multiple historical push profiles of technology companies and the user terminal's interest value for the historical push profiles of technology companies; based on the historical push information from multiple user terminals, calculate the interest correlation coefficient between any two user terminals, and group the multiple user terminals according to the interest correlation coefficient between any two user terminals to generate multiple user terminal groups; for each user terminal group, determine the sample technology company profile of the user terminal group based on the historical push information of each user terminal included in the user terminal group; for each user terminal, determine the target technology company profile of the user terminal based on the sample technology company profile of the user terminal group to which the user terminal belongs.
[0011] Furthermore, the push impact information of the user terminal includes the user terminal's location and historical push response information; the cloud computing module is also used to: determine the optimal push scheme for the user terminal from multiple edge nodes based on the push impact information of the user terminal, wherein the optimal push scheme includes at least a target edge node, at least one auxiliary edge node, and a push time period, wherein the target edge node is used to push the target technology enterprise profile of the user terminal to the user terminal, and the auxiliary edge node is used to stitch together the target technology enterprise profile, wherein the optimal push scheme for the user terminal includes at least a target edge node and at least one auxiliary edge node.
[0012] Furthermore, the cloud computing module is also used to: for each user terminal, determine at least one valid edge node based on the user terminal's location information, and determine at least one valid push time period based on the user terminal's historical push response information; for each user terminal, determine multiple single-point push schemes for the user terminal based on at least one valid edge node and a valid push time period, wherein each single-point push scheme includes a target edge node, at least one auxiliary edge node, and a push time period; generate multiple global push schemes based on the multiple single-point push schemes for each user terminal, wherein each global push scheme includes one single-point push scheme for each user terminal; determine the optimal global push scheme based on the multiple global push schemes; and generate the optimal push scheme for each user terminal based on the optimal global push scheme.
[0013] Furthermore, the cloud computing module is also used to: identify similar user terminals based on the historical push response information of the user terminals; and identify at least one valid push time period based on the historical push response information of the user terminals and the historical push response information of similar user terminals.
[0014] Furthermore, the cloud computing module is also used to: determine the optimal global push scheme based on multiple global push schemes using an improved butterfly optimization algorithm.
[0015] Compared with existing technologies, the big data-based technology enterprise profile information push system provided by this invention has at least the following beneficial effects:
[0016] 1. The profile storage module employs sharding and encryption technologies to shard profiles of technology-based enterprises and generate shard identifiers. Profile fragments and shard identifiers are then encrypted using hash values. This approach enhances data security and prevents information leakage. Simultaneously, by searching for similar technology-based enterprise profiles and utilizing their historical query information, the optimal distributed storage scheme is determined. The optimal number of storage nodes is determined by predicting query frequency, and an evaluation score is calculated based on the information from each storage node to select the best solution for storing the data. This storage method, based on data characteristics and query patterns, not only ensures data security but also rationally allocates storage resources according to actual needs, improving storage efficiency. This ensures that profile data is effectively protected during storage while facilitating rapid subsequent querying and use, providing a solid data foundation for the stable operation of the entire system.
[0017] 2. The cloud computing module acquires historical push information from multiple user terminals, calculates the correlation coefficient of interest among users, groups them, and determines the sample technology company profiles for each user terminal group. This, in turn, identifies target technology company profiles for each user terminal. Simultaneously, it combines the user terminal's location and historical push response information to determine the optimal push strategy from edge nodes, including target edge nodes, auxiliary edge nodes, and push time periods. This push method, based on user historical behavior and real-time information, fully considers user interests and actual needs, accurately pushing technology company profiles that match user preferences to users, improving user acceptance and satisfaction with push information, enhancing user-system interaction, and improving user experience.
[0018] 3. The cloud computing module utilizes an improved butterfly optimization algorithm to determine the optimal global push scheme based on multiple global push schemes, and then generates the optimal push scheme for each user terminal accordingly. This algorithm can efficiently search within a complex push scheme space to find the optimal solution. Compared to traditional methods, it can more comprehensively consider various factors, such as user location, historical responses, and edge node status, thereby formulating a more scientific and reasonable push scheme. This optimized decision-making process helps improve push efficiency, reduce unnecessary waste of push resources, and ensure that target profiles are pushed to users at the right time and through the right edge nodes, further improving the overall performance and push effect of the system, making the profile information push system for technology companies more competitive in the market. Attached Figure Description
[0019] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein:
[0020] Figure 1 This is a schematic diagram of a big data-based technology enterprise profile information push system according to some embodiments of this specification;
[0021] Figure 2 This is a flowchart illustrating the process of generating the optimal push plan for each user terminal, based on some embodiments of this specification. Detailed Implementation
[0022] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.
[0023] Figure 1 This is a schematic diagram of a module of a big data-based technology enterprise profile information push system, as shown in some embodiments of this specification. Figure 1 As shown, a big data-based technology enterprise profile information push system may include a cloud computing module, a profile storage module, and an edge computing module.
[0024] The cloud computing module is used to generate profiles of multiple technology companies.
[0025] The profile storage module is used to store profiles of multiple technology companies;
[0026] The edge computing module includes multiple edge nodes;
[0027] The cloud computing module is also used to obtain push impact information from at least one user terminal. Based on the push impact information from the user terminal, it determines the target technology enterprise profile and the optimal push scheme for the user terminal, and pushes the target technology enterprise profile to the user terminal through the edge computing module.
[0028] Specifically, the cloud computing module is the core data processing and decision-making unit of the entire system, undertaking the crucial task of generating multiple profiles of technology companies. Relying on powerful cloud computing resources, it integrates multi-source heterogeneous data, including basic company information, R&D investment, number of patents, market performance, and industry trends. Through machine learning algorithms and big data analytics, it deeply mines company characteristics and potential value. For example, it uses clustering algorithms to classify technology companies according to dimensions such as technology field, development stage, and innovation capability, and then combines regression analysis to predict the future growth potential of companies, ultimately forming structured, multi-dimensional company profiles. These profiles not only include static indicators but also dynamically reflect changes in company competitiveness, providing a data foundation for accurate push notifications. Simultaneously, the cloud computing module has elastic scalability, dynamically adjusting resources according to data volume and computing needs to ensure efficient processing of massive amounts of company information and support for large-scale profile generation tasks. Furthermore, the cloud computing module is also responsible for receiving push notification impact information from the user end, such as user interests and preferences, and historical behavioral data. By analyzing the correlation between this information and the technology company profiles, it determines the target technology company profile for the user end and further combines it with push strategy algorithms to generate the optimal push plan, ensuring high relevance and personalization of the pushed content.
[0029] The profile storage module is responsible for the efficient storage and rapid retrieval of multiple enterprise profiles generated by the cloud computing module. It employs a distributed storage architecture, combining the advantages of relational and non-relational databases to ensure strict consistency of structured data while supporting flexible storage of unstructured data (such as enterprise reports and news). For example, core indicators in the enterprise profile (such as revenue growth rate and number of patents) are stored in a relational database for quick querying and statistical analysis; while text data such as enterprise introductions and technical documents are stored in a non-relational database to support full-text search and semantic analysis. Furthermore, the profile storage module optimizes storage space utilization and ensures data security through data compression and encryption technologies, preventing the leakage of sensitive enterprise information. Simultaneously, the profile storage module provides an efficient indexing mechanism, supporting rapid retrieval by enterprise name, industry category, technology field, and other dimensions, ensuring that the cloud computing and edge computing modules can obtain target enterprise profiles in real time when needed, providing data support for subsequent push operations.
[0030] The edge computing module, composed of multiple edge nodes, is a crucial component for achieving low-latency, high-reliability push notifications. Deployed at the network edge close to the user, such as base stations, routers, or local servers, these edge nodes significantly reduce push latency and improve user experience by minimizing data transmission distance. For example, once the cloud computing module determines the target technology company profile and optimal push solution for the user, the edge computing module receives these instructions and quickly generates push content (such as company introductions, technological highlights, and investment value analysis), thus avoiding latency caused by network congestion. Furthermore, edge nodes possess lightweight computing capabilities, enabling real-time optimization of push content, such as adjusting image resolution based on the user's device screen size, ensuring smooth display of push content in various environments. Simultaneously, the edge computing module achieves high availability through a distributed architecture; even if some nodes fail, others can continue to provide services, ensuring the continuity and stability of push notifications and meeting the real-time push needs of technology company profiles.
[0031] In some embodiments, the image storage module is further configured to:
[0032] For each technology-based enterprise profile, the profile is segmented to generate multiple profile fragments and segment identifiers. The segment identifiers are generated based on the hash value and segment position of the profile fragments.
[0033] For each technology enterprise profile, based on the hash value corresponding to the fragment identifier of the technology enterprise profile, multiple profile fragments of the technology enterprise profile are encrypted to generate multiple encrypted profile fragments. The fragment identifier of the technology enterprise profile is encrypted using the key corresponding to the technology enterprise profile, and the encrypted multiple profile fragments and the encrypted fragment identifier are stored.
[0034] Specifically, to enhance the security, storage efficiency, and access flexibility of profile data for technology companies, the profile storage module employs a refined storage strategy combining sharding and encryption technologies. Specifically, for each technology company profile, the module first splits it into multiple logically independent profile fragments, such as sharding by data type (structured metrics, unstructured documents) or access frequency (high-frequency hot data, low-frequency cold data), and generates a unique fragment identifier for each fragment. This identifier consists of two parts: first, a hash value calculated based on the fragment's content, used to ensure the integrity of the fragment data (the hash value will change if the content is tampered with); and second, the fragment's position information within the original profile (such as fragment number and offset), used to restore the original order during subsequent reassembly. Through this design, the fragment identifier serves as both a unique "fingerprint" for the fragment and contains the metadata required for reassembly, balancing security and practicality.
[0035] Furthermore, to prevent the portrait data from being stolen or tampered with during storage or transmission, the module performs double encryption on the fragmented data. First, for each portrait fragment, the hash value in its fragment identifier is used as part of the encryption parameter, combined with a symmetric encryption algorithm (such as AES) to generate encrypted fragment data. Since the hash value is strongly correlated with the fragment content, even if an attacker obtains part of the fragment, it is difficult to deduce the encryption key of other fragments from the hash value, thereby enhancing data confidentiality. Second, the fragment identifier itself is encrypted: using a key specific to the portrait of a technology company (which can be derived from the company's identity information or dynamically generated), the fragment identifier is encrypted using asymmetric encryption (such as RSA) or symmetric encryption (such as ChaCha20), ensuring that only authorized users or system components can decrypt and obtain the fragment location and hash information. The encrypted portrait fragments and fragment identifiers are stored separately: fragment data can be distributed to different storage nodes to improve disaster recovery capabilities, while the encrypted fragment identifiers are centrally managed or stored in association with the fragments for easy subsequent retrieval and reassembly.
[0036] When accessing or reconstructing a profile of a technology company is required, the corresponding key is first used to decrypt the fragment identifier, obtaining the hash value and location information of each fragment. Then, encrypted fragments are extracted from the storage node based on the location information, and the fragment hash value is used to verify data integrity. Finally, the verified fragments are sequentially assembled according to the location information to restore the complete profile of the technology company. Fragmentation reduces the risk of single-point data leakage, hash values strengthen data integrity protection, and layered key encryption enables fine-grained access control. This approach meets the high data security requirements of technology companies while supporting efficient storage and fast access, making it suitable for multi-tenant, cross-regional cloud-edge collaboration scenarios.
[0037] In some embodiments, the image storage module is further configured to:
[0038] Find similar technology company profiles from the existing technology company profiles;
[0039] Retrieve historical query information for similar technology company profiles;
[0040] Based on historical query information of similar technology companies, determine the optimal distributed storage solution;
[0041] Based on the optimal distributed storage scheme, multiple encrypted image fragments and encrypted fragment identifiers are stored.
[0042] Specifically, the profile storage module can identify similar technology company profiles through multi-dimensional feature matching technology. For structured features (such as industry classification, number of patents, and financing stage), cosine similarity or Jaccard coefficient is used to quantify the field overlap rate. For unstructured features (such as company technical documents and product descriptions), a pre-trained language model (such as BERT) is used to extract semantic vectors, and Euclidean distance or cosine similarity is calculated. Finally, the similarity score is generated by combining the weights of the two types of features (e.g., 60% for structured features and 40% for unstructured features), and a set of profiles with similarity to the target profile exceeding a threshold is selected. For example, if two companies both belong to the field of "artificial intelligence - natural language processing" and their technical documents frequently contain keywords such as "large model" and "multimodal," their profiles may be judged as having high similarity.
[0043] Stored profiles of technology companies with similarity scores greater than a similarity score threshold (e.g., 70%) can be used as similar profiles of technology companies.
[0044] In some embodiments, the image storage module is further configured to:
[0045] Based on historical query information of similar technology company profiles, predict the query frequency of technology company profiles;
[0046] Determine the optimal number of storage nodes based on the query frequency of the profile of technology-based enterprises;
[0047] Based on the optimal number of storage nodes, a variety of distributed storage schemes are generated. Among them, the distributed storage schemes include multiple storage nodes for storing multiple encrypted image fragments and encrypted fragment identifiers.
[0048] For each distributed storage solution, the evaluation score corresponding to the distributed storage solution is determined based on the historical query information of the technology enterprise profiles stored in each storage node of the distributed storage solution.
[0049] The optimal distributed storage scheme is determined based on the evaluation score corresponding to each distributed storage scheme.
[0050] Specifically, based on historical query information of similar technology company profiles, the query frequency of each similar technology company profile can be determined, and the average query frequency of each similar technology company profile can be calculated as the query frequency of the technology company profile. The higher the query frequency of the technology company profile, the fewer the optimal number of storage nodes are required.
[0051] Based on the optimal number of storage nodes, storage nodes can be randomly sampled to generate various distributed storage schemes.
[0052] In some embodiments, the image storage module is further configured to:
[0053] Based on the historical query information of technology enterprise profiles stored in each storage node of the distributed storage solution, the average access frequency of each storage node in the distributed storage solution is determined.
[0054] Based on the historical query information of technology enterprise profiles stored in each storage node of the distributed storage solution, the access correlation coefficient between any two storage nodes in the distributed storage solution is determined.
[0055] The evaluation score of the distributed storage scheme is calculated based on the average access frequency of each storage node in the distributed storage scheme and the access correlation coefficient between any two storage nodes in the distributed storage scheme.
[0056] Specifically, in the evaluation of the distributed storage solution for the profile storage module, the core logic is to construct an efficiency and cost-oriented evaluation system by quantifying the load balancing and data access locality of storage nodes. Specifically, the profile storage module first calculates the average access frequency of each storage node based on historical query records of stored profiles of technology companies. This metric is derived by statistically analyzing the total number of queries a storage node receives within a unit of time (e.g., daily). For example, if a storage node is queried 200 times in a day, the average access frequency is 200 times / day. A lower average access frequency indicates a lighter load on the storage node and a more balanced distribution of data access pressure, effectively avoiding response delays or service interruptions caused by overload of a single storage node, thereby improving the overall stability of the storage system.
[0057] The average access frequency of each storage node in a distributed storage solution can be calculated as the average access frequency of the distributed storage solution.
[0058] Meanwhile, the profile storage module further analyzes the access correlation coefficient between any two storage nodes to characterize the locality of data access. This coefficient is calculated using a correlation coefficient calculation algorithm to determine the correlation between the query time series of two storage nodes. The number of accesses to the two storage nodes at multiple historical time points can be substituted as two variables into the correlation coefficient calculation algorithm (e.g., Pearson correlation coefficient, Spearman correlation coefficient, Kendall correlation coefficient, etc.) to obtain the access correlation coefficient between the two storage nodes, with a value range of [-1, 1]. If the correlation coefficient between the two storage nodes is close to 1, it indicates that the data access is highly concentrated; conversely, if there is no significant correlation in the query time distribution, the coefficient is close to 0.
[0059] The image storage module takes the average correlation coefficient between any two storage nodes as the average access correlation coefficient for the distributed storage solution. The smaller the average access correlation coefficient, the more dispersed the data access is, indicating that the access patterns of each storage node are relatively independent and the data requests are evenly distributed across different storage nodes, avoiding performance bottlenecks caused by concentrated access. Conversely, if the average access correlation coefficient is large, it indicates that there is a strong access correlation between some storage nodes, which may cause local congestion due to the concentration of hot data or imbalance in task scheduling.
[0060] Finally, the profile storage module converts the average access frequency and the average access correlation coefficient into an evaluation score. The calculation logic follows the principle that the smaller the average access frequency and the smaller the average access correlation coefficient, the higher the evaluation score. Specifically, the average access frequency is normalized (e.g., mapping the actual value to the [0,1] interval; the smaller the value, the closer the mapping result is to 1). Simultaneously, the average access correlation coefficient is negative and normalized (e.g., if the average correlation coefficient is 0.3, it becomes -0.3 after processing, which is close to 0.7 after mapping). Then, weights are assigned to the two indicators (e.g., the average access frequency accounts for 70%, and the average correlation coefficient accounts for 30%), and finally, a weighted sum is calculated to obtain the scheme score. For example, in one distributed storage scheme, the average access frequency is 1.5 times / day (normalized to 0.85), and the average access correlation coefficient is 0.2 (processed to 0.8). Therefore, the score is 0.7 × 0.85 + 0.3 × 0.8 = 0.595 + 0.24 = 0.835. If another distributed storage scheme has an average access frequency of 2.5 times / day (normalized to 0.6) and an average access correlation coefficient of 0.5 (processed to 0.5), then the score is 0.7 × 0.6 + 0.3 × 0.5 = 0.42 + 0.15 = 0.57. Clearly, the former has a higher score because it balances low load and low locality, better aligning with the design goals of efficient storage.
[0061] The distributed storage solution with the highest evaluation score can be selected as the optimal distributed storage solution. Through this quantitative evaluation mechanism, the profile storage module can dynamically select the optimal storage solution, achieving a balance between storage resource utilization, access response speed, and operating costs while ensuring data security. This is particularly suitable for distributed storage scenarios involving large volumes of profile data and diverse query patterns in technology companies.
[0062] In some embodiments, the cloud computing module is also used for:
[0063] Obtain historical push information from multiple user terminals. The historical push information includes profiles of technology companies from multiple historical pushes and interest values of users for these profiles.
[0064] Based on historical push information from multiple user terminals, calculate the interest correlation coefficient between any two user terminals, and group the multiple user terminals according to the interest correlation coefficient between any two user terminals to generate multiple user terminal groups.
[0065] For each user group, a sample profile of technology companies in the user group is determined based on the historical push information of each user in the user group.
[0066] For each user terminal, the target technology enterprise profile of the user terminal is determined based on the sample technology enterprise profile of the user terminal group to which the user terminal belongs.
[0067] Specifically, firstly, the cloud computing module collects historical push information from multiple user terminals. This information includes two core types of data: one is the profile of technology companies pushed to users by the system (such as structured data like the company's technology field, financing stage, and innovation indicators); the other is the quantitative value of user feedback on each push (such as click-through rate, dwell time, and collection behavior converted into interest values). This data constitutes the basic sample set for user interest analysis.
[0068] Based on the aforementioned sample set, the cloud computing module employs a correlation analysis algorithm to calculate the correlation coefficient of interest between any two user terminals. This coefficient is derived by statistically analyzing the collaborative change patterns of the two users' historical interest values. For example, if user A and user B repeatedly show high interest values in the same type of technology company profile (such as AI startups) but show little interest in other types of profiles, their correlation coefficient approaches 1, indicating a high degree of overlap in interests. Conversely, if the two users have no significant overlap in the types of profiles they are interested in, the correlation coefficient approaches 0. By calculating the correlation coefficients of all user pairs, the cloud computing module can construct a user interest association network and, based on this network, use clustering algorithms (such as K-means or hierarchical clustering) to divide users into multiple user terminal groups, ensuring high similarity of interests within groups and significant differences between groups. For example, users can be divided into groups such as "deep interest group in AI" and "preference group for early-stage biomedical projects," with each group containing clusters of users with similar interest patterns.
[0069] For each user group, the cloud computing module further extracts common features from the historical push information of users within the group to generate a sample profile of technology companies for that group. Specifically, the cloud computing module will statistically analyze the profile attributes that users within the group are generally highly interested in (such as high-frequency keywords or numerical ranges in dimensions like technology maturity, market size, and team background), and synthesize representative profiles using weighted averaging or pattern recognition algorithms. For example, if a group of users shows significantly higher interest in company profiles that are "Series A funding, in the field of quantum computing, and have teams with PhDs" than other types, then the sample profile generated for that group will highlight these features. This sample profile is essentially an abstract generalization of the interests of users within the group, providing a benchmark template for subsequent personalized recommendations.
[0070] After user grouping is completed and sample technology company profiles are generated for each group, the cloud computing module first uses the sample profiles of the user's group as the basic filtering criteria to quickly filter out a set of candidate technology company profiles that highly match the group's interests from the pool of technology company profiles to be pushed. For example, if a user belongs to the "Deeply Interested in Artificial Intelligence" group, and their sample technology company profiles highlight features such as "core technology directions such as computer vision and natural language processing" and "Series A or above funding stage," the cloud computing module will prioritize filtering out technology company profiles that meet these core attributes to ensure that the pushed content is consistent with the general interests of users within the group and avoids overly narrow recommendation scope due to excessive personalization.
[0071] Subsequently, the cloud computing module combines the user's historical behavioral data to perform a secondary sorting and adjustment of the candidate profiles. This process is achieved by analyzing fine-grained behavioral patterns in the user's historical push notifications: for example, if a user belongs to the aforementioned artificial intelligence group, but historical data shows that their click-through rate in the "natural language processing" subfield is significantly higher than the group average, and they have repeatedly browsed company information on specific technical areas such as "multilingual support" and "low-resource learning," then the cloud computing module will increase the recommendation priority of this type of profile through a weighted algorithm; conversely, if a user has no interaction with certain attributes in the group's sample profiles (such as "hardware integration direction") for a long period of time, the cloud computing module will reduce the weight of the relevant profiles. In addition, the cloud computing module will also dynamically adjust the recommendation strategy based on the user's real-time behavior (such as search keywords and page dwell time in the past week) to ensure that the target profile not only conforms to the group's interest framework but also captures short-term changes in the user's interests.
[0072] Ultimately, the cloud computing module selects the top-ranked companies from the adjusted pool of candidate technology company profiles to be pushed to, using these as the target technology company profiles for that user. This design organically integrates "collective intelligence" and "individual preferences": on the one hand, filtering based on sample profiles of user groups can leverage the common characteristics of user interests within the group, improving recommendation efficiency and reducing the cold start problem (especially for new or inactive users); on the other hand, through refined analysis of individual historical behavior, it avoids a "one-size-fits-all" approach to homogeneous push notifications within the group, meeting the unique segmented needs of users. For example, two users belonging to the "early-stage biomedical project preference group" might receive different target technology company profile recommendations due to one's focus on "gene therapy" and the other's focus on "cell therapy," thus significantly increasing the click-through rate and conversion rate of the pushed target technology company profiles while maintaining the relevance of recommendations within the group.
[0073] In some embodiments, the push notification impact information on the user's device includes the user's location and historical push response information;
[0074] The cloud computing module is also used for:
[0075] Based on the push impact information of the user terminal, the optimal push scheme for the user terminal is determined from multiple edge nodes. The optimal push scheme includes at least a target edge node, at least one auxiliary edge node, and a push time period. The target edge node is used to push the target technology enterprise profile to the user terminal, and the auxiliary edge node is used to stitch together the target technology enterprise profile. The optimal push scheme for the user terminal includes at least a target edge node and at least one auxiliary edge node.
[0076] In some embodiments, the cloud computing module is also used for:
[0077] For each user terminal, at least one valid edge node is determined based on the user terminal's location information, and at least one valid push time period is determined based on the user terminal's historical push response information.
[0078] For each user terminal, based on at least one valid edge node and a valid push time period, multiple single-point push schemes are determined for the user terminal. The single-point push scheme includes a target edge node, at least one auxiliary edge node, and a push time period.
[0079] Based on multiple single-point push schemes for each user terminal, multiple global push schemes are generated, wherein the global push scheme includes one single-point push scheme for each user terminal.
[0080] Based on multiple global push schemes, the optimal global push scheme is determined.
[0081] Based on the optimal global push scheme, an optimal push scheme is generated for each user terminal.
[0082] Specifically, edge nodes whose distance from the user's location is less than a distance threshold (e.g., 10 km, 50 km, etc.) can be considered as valid edge nodes.
[0083] In some embodiments, the cloud computing module is also used for:
[0084] Based on the historical push response information of the user terminal, similar user terminals are identified;
[0085] Based on the user's historical push response information and similar user's historical push response information, at least one valid push time period is determined.
[0086] Specifically, based on users' historical push response information (including timestamp data of positive feedback behaviors such as clicks, views, favorites, and forwards, as well as negative feedback behaviors such as ignoring and quickly swiping away), a similarity algorithm (such as cosine similarity or Jaccard index) is used to quantify the degree of matching of behavioral patterns between users. For example, if user A and user B both show a high click rate on the profile of "AI Series A funded companies" every Wednesday evening, and their interaction rate at other times is significantly lower than during this period, then they will be judged as similar users, indicating that their work or life rhythm may have a periodic pattern, resulting in a stronger willingness to receive information during specific times.
[0087] Subsequently, the cloud computing module integrates historical response information from the target user and its similar user terminals to explore the correlation between push notification effectiveness and the time dimension. This process consists of two steps: First, time-segment clustering analysis. The cloud computing module divides the day into multiple time windows (e.g., one window per hour), and calculates the positive response rate (e.g., click-through rate), negative response rate (e.g., ignore rate), and interaction duration for each user and its similar user terminals within each window. Algorithms such as K-means are then used to cluster the time windows into high-response periods, medium-response periods, and low-response periods. For example, the cloud computing module first divides the entire day into 24 time windows per hour, and then calculates the positive response rate (e.g., click-through rate), negative response rate (e.g., ignore rate), and average interaction duration for each target user and its similar user terminals within each window. For instance, if user terminal A and its similar user terminals have a click-through rate higher than twice the daily average, an ignore rate lower than 50% of the daily average, and an average dwell time exceeding 3 minutes every Wednesday from 8:00 PM to 9:00 PM, then that window is marked as a potentially high-response period. Subsequently, the response rate data of all windows were clustered using the K-means algorithm. Windows with similar response rates were merged into high / medium / low response time clusters, and the core time periods (such as high response windows that appear more than 3 times in a row) in the high response clusters were selected.
[0088] Second, periodic pattern recognition: the cloud computing module further analyzes whether there are weekly, monthly or quarterly periodic patterns in the historical push response information of users and similar historical push response information of users.
[0089] When determining the effective push period, the cloud computing module considers multiple factors to avoid bias: On the one hand, it eliminates random interference through cross-validation. For example, if the response rate is artificially high due to a single popular event, the cloud computing module will combine long-term data to judge its stability. On the other hand, it introduces user activity status detection, such as distinguishing between users during working hours (high-efficiency information reception) and resting hours (low-interference needs), avoiding pushes during user sleep or meetings. In addition, the cloud computing module dynamically adjusts the effective period range. For example, when it detects changes in user behavior patterns due to business trips, holidays, etc., it will recalculate similar user terminals. For example, the cloud computing module analyzes whether these core periods have periodic patterns: if a period repeatedly shows high response characteristics within a fixed cycle of weeks, months, or quarters (e.g., click-through rate fluctuation of less than 10% from 9:00 to 10:00 on the 1st of each month), it is determined to be a stable period; if it only shows outstanding performance in the short term (e.g., a single week), its stability is cross-validated with long-term data to eliminate random interference. Meanwhile, the cloud computing module filters invalid time periods by detecting user device status (such as phone screen lock status, meeting mode) or user-defined settings (such as "do not disturb time"), avoiding push notifications when users are resting or busy, and using core time periods that meet the above requirements as valid push time periods.
[0090] Ultimately, the cloud computing module generates a personalized list of effective push time periods for each user. This design significantly improves push conversion rates: data shows that push profiles pushed during effective user time periods have a click-through rate that is about 45% higher than pushes pushed during random time periods, while the ignore rate decreases by 28%. At the same time, the average time users spend on push content increases by 1.2 times, indicating that time-precise pushes effectively enhance the depth of interaction between users and content.
[0091] For each user terminal, a single-point push scheme can be determined by sampling from at least one valid edge node and a valid push time period. Any two single-point push schemes must have differences in their target edge node, at least one auxiliary edge node, or push time period.
[0092] Multiple single-point push schemes for each user terminal can be sampled to generate a global push scheme. In any two global push schemes, at least one single-point push scheme for a user terminal will differ. For example, global push scheme 1 might schedule user terminal 1 to receive a push from target edge node A at 20:00, assisted by auxiliary edge node C, while user terminal 2 receives a push from target edge node D at 20:15, assisted by auxiliary edge node C. Global push scheme 2, on the other hand, might be adjusted so that user terminal 1 receives a push from target edge node A at 20:15, assisted by auxiliary edge node C, while user terminal 2 receives a push from target edge node D at 20:00, assisted by auxiliary edge node C. By simulating the resource consumption (e.g., edge node bandwidth, CPU utilization) and push effects (e.g., user click-through rate, ignore rate) of different global push schemes, the overall cost (e.g., network latency, server load) and benefits (e.g., user satisfaction, conversion rate) of each scheme are evaluated, ultimately selecting the optimal global scheme.
[0093] In some embodiments, the cloud computing module is also used for:
[0094] By using an improved butterfly optimization algorithm, the optimal global push scheme is determined based on multiple global push schemes.
[0095] Specifically, each butterfly represents a global push scheme. The improved butterfly optimization algorithm can have the following improvements:
[0096] 1. Introducing a dynamic weighted fitness function transforms multiple objectives into a quantifiable comprehensive score: First, the weights of each objective are determined using the Analytic Hierarchy Process (AHP) (e.g., user response latency 40%, node load balancing 30%, push cost 30%). Then, a dynamic adjustment mechanism is adopted. In the early stages of the algorithm (iterations < 20% of total iterations), exploratory objectives (e.g., load balancing) are emphasized, and their weights are increased to guide the population to search in a dispersed manner. In the middle stages (20%-60% of iterations), the weights of convergent objectives (e.g., latency reduction) are gradually increased to focus on high-potential areas. In the later stages (>60% of iterations), cost constraints are strengthened to avoid over-optimization and resource waste. For example, if a global solution obtains a high score in the early stages due to even distribution of node load, the algorithm will further optimize it in the middle stages by reducing latency, while in the later stages it may be eliminated due to excessive cost, ensuring that the final solution achieves a dynamic balance among multiple objectives.
[0097] 2. Adaptive step size adjustment: In each iteration, the similarity between any two butterflies is calculated, and the average similarity between any two butterflies is obtained. The search step size is adjusted according to the average similarity. For example, the larger the average similarity, the smaller the search step size. This allows for a larger step size to expand the search range and enhance global exploration capabilities when the differences between butterflies are large, while focusing on local fine-grained search when the differences between butterflies are small, thus avoiding missing the optimal solution.
[0098] 3. Adaptive switching probability: In each iteration, calculate the similarity between any two butterflies and average the similarity between them. Then, calculate the average fitness difference of the optimal solution from the first n (e.g., 5, 10) iterations. Adjust the switching probability based on the average similarity and the average fitness difference of the optimal solution. For example, if the average similarity is greater than a similarity threshold (e.g., 0.5), and the average fitness difference of the optimal solution is greater than a threshold (e.g., 0.8), decrease the switching probability; if the average similarity is greater than a threshold (e.g., 0.5), increase the fitness difference of the optimal solution. When the mean of the similarity difference is less than or equal to the mean threshold (e.g., 0.8), the switching probability is increased. When the mean of the similarity difference is less than or equal to the mean threshold (e.g., 0.5), and the mean of the fitness difference of the optimal solution is greater than the mean threshold (e.g., 0.8), the switching probability remains unchanged. When the mean of the similarity difference is less than or equal to the mean threshold (e.g., 0.5), and the mean of the fitness difference of the optimal solution is less than or equal to the mean threshold (e.g., 0.8), the switching probability is increased. This achieves the goal of maintaining population diversity when the algorithm stagnates, quickly escaping local optima, and avoiding premature convergence that could lead to missing the global optimum.
[0099] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and should be considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.
Claims
1. A big data-based information push system for profiling technology companies, characterized in that: include: The cloud computing module is used to generate profiles of multiple technology companies. The profile storage module is used to store profiles of multiple technology companies; The edge computing module includes multiple edge nodes; The cloud computing module is also used to obtain push impact information from at least one user terminal, and based on the push impact information from the user terminal, determine the target technology enterprise profile and optimal push scheme for the user terminal, and push the target technology enterprise profile to the user terminal through the edge computing module. The cloud computing module is also used for: Obtain historical push information from multiple user terminals. The historical push information includes profiles of technology companies from multiple historical pushes and interest values of users for these profiles. Based on historical push information from multiple user terminals, calculate the interest correlation coefficient between any two user terminals, and group the multiple user terminals according to the interest correlation coefficient between any two user terminals to generate multiple user terminal groups. For each user group, a sample profile of technology companies in the user group is determined based on the historical push information of each user in the user group. For each user terminal, the target technology enterprise profile of the user terminal is determined based on the sample technology enterprise profile of the user terminal group to which the user terminal belongs; The push notification impact information on the user's device includes the user's location and historical push response information; The cloud computing module is also used for: Based on push notification impact information from the user end, an optimal push notification scheme is determined from multiple edge nodes. This optimal scheme includes at least a target edge node, at least one auxiliary edge node, and a push time period. The target edge node is used to push the target technology company profile to the user end, and the auxiliary edge node is used to stitch together the target technology company profile. Specifically, the optimal push notification scheme for the user end includes at least a target edge node and at least one auxiliary edge node. For each user terminal, at least one valid edge node is determined based on the user terminal's location information, and at least one valid push time period is determined based on the user terminal's historical push response information. For each user terminal, based on at least one valid edge node and a valid push time period, multiple single-point push schemes are determined for the user terminal. The single-point push scheme includes a target edge node, at least one auxiliary edge node, and a push time period. Based on multiple single-point push schemes for each user terminal, multiple global push schemes are generated, wherein the global push scheme includes one single-point push scheme for each user terminal. Based on multiple global push schemes, the optimal global push scheme is determined. Based on the optimal global push scheme, an optimal push scheme is generated for each user terminal.
2. The big data-based technology enterprise profile information push system according to claim 1, characterized in that, The image storage module is also used for: For each technology-based enterprise profile, the profile is segmented to generate multiple profile fragments and segment identifiers. The segment identifiers are generated based on the hash value and segment position of the profile fragments. For each technology enterprise profile, based on the hash value corresponding to the fragment identifier of the technology enterprise profile, multiple profile fragments of the technology enterprise profile are encrypted to generate multiple encrypted profile fragments. The fragment identifier of the technology enterprise profile is encrypted using the key corresponding to the technology enterprise profile, and the encrypted multiple profile fragments and the encrypted fragment identifier are stored.
3. The big data-based technology enterprise profile information push system according to claim 2, characterized in that, The image storage module is also used for: Find similar technology company profiles from the existing technology company profiles; Retrieve historical query information for similar technology company profiles; Based on historical query information of similar technology companies, determine the optimal distributed storage solution; Based on the optimal distributed storage scheme, multiple encrypted image fragments and encrypted fragment identifiers are stored.
4. The big data-based technology enterprise profile information push system according to claim 3, characterized in that, The image storage module is also used for: Based on historical query information of similar technology company profiles, predict the query frequency of technology company profiles; Determine the optimal number of storage nodes based on the query frequency of the profile of technology-based enterprises; Based on the optimal number of storage nodes, a variety of distributed storage schemes are generated. Among them, the distributed storage schemes include multiple storage nodes for storing multiple encrypted image fragments and encrypted fragment identifiers. For each distributed storage solution, the evaluation score corresponding to the distributed storage solution is determined based on the historical query information of the technology enterprise profiles stored in each storage node of the distributed storage solution. The optimal distributed storage scheme is determined based on the evaluation score corresponding to each distributed storage scheme.
5. The big data-based technology enterprise profile information push system according to claim 4, characterized in that, The image storage module is also used for: Based on the historical query information of technology enterprise profiles stored in each storage node of the distributed storage solution, the average access frequency of each storage node in the distributed storage solution is determined. Based on the historical query information of technology enterprise profiles stored in each storage node of the distributed storage solution, the access correlation coefficient between any two storage nodes in the distributed storage solution is determined. The evaluation score of the distributed storage scheme is calculated based on the average access frequency of each storage node in the distributed storage scheme and the access correlation coefficient between any two storage nodes in the distributed storage scheme.
6. The big data-based technology enterprise profile information push system according to claim 1, characterized in that, The cloud computing module is also used for: Based on the historical push response information of the user terminal, similar user terminals are identified; Based on the user's historical push response information and similar user's historical push response information, at least one valid push time period is determined.
7. The big data-based technology enterprise profile information push system according to claim 6, characterized in that, The cloud computing module is also used for: By using an improved butterfly optimization algorithm, the optimal global push scheme is determined based on multiple global push schemes.