Data management method and device, computer equipment and storage medium
By integrating multi-dimensional data and calculating industry benchmarks, combined with intelligent analysis of large language models, the problem of data silos in fund marketing and operations has been solved, enabling cross-channel data sharing and refined operations, thereby improving operational efficiency and decision support capabilities.
Patent Information
- Application Number
- CN202511673732.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-14
- Publication Date
- 2026-02-13
AI Technical Summary
Existing technological solutions cannot effectively solve the data silo problem in fund marketing and operations. They lack the ability to deeply process data, make horizontal comparisons with other industries, and provide intelligent decision support, resulting in low operational efficiency and limited autonomy.
By acquiring and integrating multi-dimensional data, we obtain comprehensive industry data and perform benchmark calculations. We then use large language models for intelligent analysis to generate data management strategies.
Enables cross-channel and cross-platform data sharing and in-depth analysis, improves the automation and efficiency of data processing, enhances the ability to make horizontal comparisons across industries and the level of refined operation, and provides data-driven decision support.
Smart Images

Figure CN121526792A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a data management method, apparatus, computer equipment, and storage medium. Background Technology
[0002] In the current fund sales industry, fund companies generally collaborate with various sales channels to develop marketing zones within their mobile applications to promote their products. These zones typically take the form of lightweight pages based on H5 technology, WeChat mini-programs, or the "Wealth Account" feature provided by channel platforms.
[0003] Specifically, H5 page solutions are web pages developed based on HTML5 technology, which are embedded within the framework of fund sales channels (such as bank apps, third-party payment platforms, etc.) to display content. Their technical architecture mainly consists of front-end H5 page code and back-end interfaces. The front-end is responsible for page display and user interaction, while the back-end provides data interfaces. However, this solution typically lacks an independent back-end management system; content and data updates require manual modification of code or back-end interfaces. H5 pages usually include modules such as header images, market information, investment research opinions, fund product displays, marketing materials, and risk warnings. This content is relatively fixed, with a low update frequency and a cumbersome update process. In terms of data management, H5 pages typically collect data through simple page tracking, such as page views and clicks, but lack the ability to manage and share deeper data such as product transaction data and customer profile data. The advantages of H5 pages are relatively low development costs and short deployment cycles. Its disadvantages are: operation and maintenance are highly dependent on manual labor, with a low degree of automation; content updates are inflexible; data sharing and integration capabilities are weak, making it difficult to achieve refined operation; it cannot provide comprehensive industry-level data for horizontal comparison, nor can it achieve intelligent operation guidance.
[0004] A fund company's "Wealth Account" solution refers to an official account operated by the fund company on a third-party platform, similar to a dedicated brand space. Its technical architecture is based on the underlying technical framework of the third-party platform. The fund company uses the platform's provided interfaces and backend tools to publish content, manage products, and interact with customers. "Wealth Accounts" typically have rich operational functions, including article publishing, live interaction, investor education content updates, and product showcases. Operators can flexibly publish and adjust content using the platform's backend tools. Regarding data management, the platform will share some operational data with the fund company, such as content readership and interaction data. However, core transaction data and customer profile data are usually controlled by the platform, and the fund company cannot directly access and integrate them. The advantages of this solution are: a fixed user base, facilitating brand promotion and customer operations; and the platform also provides certain data analysis tools to help understand operational effectiveness. The disadvantages of this solution are: high dependence on the third-party platform, leading to limited operational autonomy; limited data sharing and interoperability capabilities, making it difficult to integrate data with other channels; and the lack of data dashboard functionality, hindering cross-channel data analysis and horizontal comparison.
[0005] The fund mini-program solution is a lightweight application developed based on a super app, allowing users to use it without downloading. Its technical architecture relies on a specific platform's development framework, possessing independent pages and functional modules. This mini-program can interface with the fund company's proprietary systems (such as trading systems and customer relationship management systems) to achieve core functions like trading and querying. Compared to H5, mini-programs can support more complex functions, including product display, order placement, fixed investment settings, and asset querying. Content and function updates are relatively flexible and can be managed through the backend. In terms of data management, the mini-program can acquire some user behavior data and interface with the fund company's internal systems through APIs to achieve partial data interoperability. However, the implementation of data sharing and security mechanisms depends on complex system integration and development. The advantages of this solution are a smooth user experience, lower development costs than native apps, the ability to implement more complex trading functions, and the ability to interface with internal company systems for data interoperability. However, its disadvantages include: limited data sharing capabilities, difficulty in achieving secure data sharing with sales channels or other platforms; typically lacking configuration middleware and data dashboards; and reliance on other systems for operations and data analysis, hindering integrated management.
[0006] In summary, the core deficiency of existing technological solutions lies in their inability to fundamentally solve the "data silo" problem, as well as their lack of capabilities for in-depth data processing, cross-industry comparison, and intelligent decision support. These shortcomings result in significant obstacles for fund marketing operations in areas such as data analysis, efficiency improvement, and intelligent application. Summary of the Invention
[0007] This invention provides a data management method, apparatus, computer equipment, and storage medium, aiming to solve the data silo problem in the prior art, improve the ability of in-depth data processing, cross-industry comparison, and intelligent decision support, thereby improving the data management effect.
[0008] In a first aspect, embodiments of the present invention provide a data management method, including: Acquire multi-dimensional data from different data sources and integrate the multi-dimensional data to obtain multi-source heterogeneous data; Acquire industry panoramic data and perform benchmark calculations on the industry panoramic data to obtain industry average data; Based on the aforementioned multi-source heterogeneous data and industry average data, intelligent analysis is performed using a large language model, and corresponding data management strategies are output.
[0009] Secondly, embodiments of the present invention provide a data management system, including: The first acquisition unit is used to acquire multi-dimensional data from different data sources and integrate the multi-dimensional data to obtain multi-source heterogeneous data. The second acquisition unit is used to acquire industry panoramic data and perform benchmark calculations on the industry panoramic data to obtain industry average data. The intelligent analysis unit is used to perform intelligent analysis based on the multi-source heterogeneous data and industry average data, using a large language model, and output corresponding data management strategies.
[0010] Thirdly, embodiments of the present invention provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the data management method as described in the first aspect.
[0011] Fourthly, embodiments of the present invention provide a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, which, when executed by a processor, implements the data management method as described in the first aspect.
[0012] This invention provides a data management method, apparatus, computer equipment, and storage medium. The method includes: acquiring multi-dimensional data from different data sources and integrating the multi-dimensional data to obtain multi-source heterogeneous data; acquiring industry panoramic data and performing benchmark calculations on the industry panoramic data to obtain industry average data; and based on the multi-source heterogeneous data and industry average data, performing intelligent analysis using a large language model and outputting corresponding data management strategies. This invention, through the integration of multi-source heterogeneous data and the calculation of industry benchmark data, achieves cross-channel and cross-platform data sharing and in-depth analysis, effectively breaking down the "data silos" phenomenon in traditional solutions. Simultaneously, utilizing the intelligent analysis capabilities of the large language model, it can generate targeted management strategies based on real-time data and industry average levels, providing data-driven decision support for fund marketing operations. This not only improves the automation and efficiency of data processing but also significantly enhances the ability for horizontal industry comparison and the level of refined operation, thereby improving data management effectiveness. Attached Figure Description
[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 A flowchart illustrating a data management method provided in an embodiment of the present invention; Figure 2 This is a schematic block diagram of a data management system provided in an embodiment of the present invention. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0017] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0018] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0019] Please see below. Figure 1 The data management method provided in this embodiment of the invention specifically includes steps S101 to S103.
[0020] Step S101: Obtain multi-dimensional data from different data sources and integrate the multi-dimensional data to obtain multi-source heterogeneous data; Step S102: Obtain industry panoramic data and perform benchmark calculations on the industry panoramic data to obtain industry average data; Step S103: Based on the multi-source heterogeneous data and industry average data, perform intelligent analysis using a large language model and output the corresponding data management strategy.
[0021] In this embodiment, raw data covering multiple dimensions such as user behavior, transaction records, and market dynamics is first collected from various data sources. This data is then integrated into a multi-source heterogeneous dataset. Next, through collaboration with industry associations and market research institutions, comprehensive industry data is obtained, and benchmark calculations are performed to obtain representative industry average data. Finally, the integrated multi-source heterogeneous data and the calculated industry average data are input into a pre-trained large language model. The large language model's natural language processing capabilities and deep learning algorithms are used to perform deep analysis and pattern recognition on the input data, generating targeted data management strategy recommendations.
[0022] This embodiment achieves comprehensive coverage and in-depth analysis of cross-channel and cross-platform data through the integration of multi-source heterogeneous data and the calculation of industry benchmark data. Compared with the drawbacks of traditional solutions where data is scattered and difficult to integrate, this embodiment effectively breaks down the "data silo" phenomenon, providing a broader perspective and a more solid basis for data management. Simultaneously, leveraging the intelligent analysis capabilities of large language models, it can dynamically compare real-time data with industry averages, accurately identify anomalies and potential opportunities in the data, and generate highly targeted and actionable management strategy recommendations. These strategy recommendations not only cover all aspects of data collection, processing, and analysis, but also involve multiple business areas such as product promotion, customer operations, and risk control, providing comprehensive, data-driven decision support for fund marketing operations. Thus, it not only significantly improves the automation and efficiency of data processing, but also greatly enhances the effectiveness of data management through industry-wide comparisons and refined operations, promoting the development of fund marketing operations towards intelligence and refinement.
[0023] In practical applications, a marketing management platform can be built based on the data management method provided in this embodiment. This platform can include a front-end page, a configuration platform, and a data dashboard. The front-end page refers to the H5 page launched on the sales channel app. This page can include integrated modules such as header images, market information, fund company investment research views, fund product displays, marketing materials, and risk warnings. This page can connect to objects such as funds or all gold-related funds in the market, and can also implement functions such as customer investment education. The configuration platform is a platform that can be operated after logging in via a website and verifying identity. Through this platform, the element layout and product list of the front-end page can be modified. The data dashboard is a platform that connects the front-end page and linked products, and presents data such as tracking data, marketing data funnels, customer profiles, customer transaction dimension data, and product sales dimension data through front-end page tracking and the Bosera Funds trading system. When developing the front-end page, you can select the corresponding modules and upload customized content to generate a gold-themed front-end page suitable for the new channel.
[0024] Operation of the configuration platform: (1) Automated updates: The configuration platform is associated with the fund company's internal content operation platform, data interface, and fund manager-related emails. The content of the corresponding modules on the front-end page of the Gold Zone will be automatically updated. (2) Product list maintenance: manually configure the product list and adjust the product display order.
[0025] Data integration of the data dashboard: (1) Data funnel (marketing efficiency data) display, including page tracking data UV, PV, purchase click rate, and query the number of new customers (above 10 yuan) and subscription data of the corresponding products. It can calculate the average number of customer visits, UV ratio, visit purchase click rate, purchase conversion rate, and average subscription order price. (2) Customer profile data, product transaction data, and data from fund companies; (3) Industry average data: the arithmetic average of data from different channels with the same caliber is used as the reference data for each independent channel.
[0026] Through the aforementioned marketing platform, the three major functional modules—front-end user experience, back-end operation management, and data intelligent analysis—can be organically integrated into a highly efficient closed-loop system. The front-end page, as the window for user interaction, boasts advantages such as short development and deployment cycles, low investment costs, and the ability to achieve matrix management. The configuration platform, as the central hub for operation and maintenance, simplifies operation and maintenance while providing automated update capabilities, reducing labor costs. The data dashboard, as the brain of decision support, enables secure sharing between fund companies and sales channels, including page tracking data, product transaction data, customer profile data, and customer transaction dimension data. These three components, through strictly defined data interfaces and communication protocols, achieve smooth data flow and collaborative operation, providing industry-wide data comparisons and, by incorporating a large language model, intelligently providing guidance for zone operation and iteration.
[0027] In one embodiment, the step of acquiring multi-dimensional data from different data sources and integrating the multi-dimensional data to obtain multi-source heterogeneous data includes: The multi-dimensional data is extracted from different data sources; wherein, the multi-dimensional data includes front-end page data, transaction system data, customer profile data, and external cooperation channel data; The multi-dimensional data is preprocessed to obtain the multi-source heterogeneous data.
[0028] The data sources described in this embodiment include front-end pages, fund company trading system data, customer profile data platforms, and external cooperation channels. This allows for the acquisition of corresponding front-end page data, trading system data, customer profile data, and external cooperation channel data. Specifically, when acquiring front-end page data, user behavior data, such as UV (unique visitors), PV (page views), page dwell time, and click-through rate, can be collected in real time by embedding tracking scripts on the front-end page. When acquiring trading system data, customer transaction data, such as subscription and redemption records, subscription amounts, and holding data, can be acquired in batches or in real time. For customer profile data, static and dynamic customer profile data can be obtained from the fund company's internal customer management system. For external cooperation channel data, industry benchmark data from other fund companies or channels can be obtained through the secure sharing mechanism detailed later.
[0029] Subsequently, the acquired multi-dimensional data is preprocessed accordingly to integrate it into multi-source heterogeneous data.
[0030] Specifically, the preprocessing of the multi-dimensional data to obtain the multi-source heterogeneous data includes: The multi-dimensional data is standardized using the Z-fraction scaling algorithm; The multi-dimensional data is normalized using the Min-Max normalization algorithm; The three-standard-deviation algorithm is used to perform outlier detection processing on the multi-dimensional data; The multi-dimensional data is then cleaned and de-identified.
[0031] This embodiment ensures consistency, comparability, and accuracy by deeply processing multi-dimensional data. First, data from different channels may have significantly different statistical definitions and units. To make data from different channels (such as purchase click-through rate and average order value) comparable, this embodiment employs two standardization techniques: (1) Z-score scaling: When the data points follow a normal distribution, it converts the data points into a multiple of the standard deviation. The formula is: x′=(x-μ) / σ; Where μ is the mean and σ is the standard deviation. Z-score scaling effectively handles outliers in data, forcing them to be scaled to a controllable range, ensuring that data of different dimensions are transformed into a comparable, unified standard.
[0032] (2) Min-Max Normalization: Also known as Min-Max Scaling, it maps the value range of original data features to a specified interval, typically [0, 1] or other custom ranges (such as [-1, 1]). It can eliminate dimensional differences between different features, preventing certain features from dominating model training due to excessively large numerical ranges. Simultaneously, it preserves the relative relationships and distribution shape of the data, but it is sensitive to outliers, which can affect min and max, leading to compression of normal data. The basic normalization formula maps data x to the interval [0, 1]: ; If mapping to a custom interval [a, b] (e.g., [-1, 1]) is required, the expanded formula is: ; Where x' is the mapped data, min represents the minimum value, and max represents the maximum value.
[0033] Secondly, during the data integration process, outliers (such as extremely high or low data due to technical errors or cheating) can severely interfere with subsequent analysis. Therefore, this embodiment applies the 3σ criterion (three standard deviations) for outlier detection. Any data point exceeding the range of the mean plus or minus three standard deviations will be identified as an outlier. These outliers will be removed or corrected according to preset business rules to ensure the accuracy of data analysis.
[0034] This embodiment also cleanses customer transaction data and profile data, removing incomplete or inconsistent records. Furthermore, to protect user privacy, all sensitive customer data undergoes rigorous anonymization and desensitization before entering the analysis phase.
[0035] In one embodiment, acquiring industry panoramic data and performing benchmark calculations on the industry panoramic data to obtain industry average data includes: A multi-party secure computation model and anonymity query algorithm are used to obtain panoramic industry data from multiple related industry data providers.
[0036] This embodiment, through a multi-party secure computation model and an anonymous query algorithm, enables the secure sharing and acquisition of comprehensive industry data while protecting the data privacy of all parties. Specifically, the multi-party secure computation model allows multiple participants to jointly complete data computation tasks without disclosing their individual original data, thereby obtaining the required comprehensive industry data. The anonymous query algorithm further ensures the anonymity of data during the query process, preventing the risk of data leakage.
[0037] In practical applications, a data sharing architecture can be built, which includes: Data providers: The participating fund companies, which will input their locally pre-processed and anonymized marketing data (such as conversion rates, average order value, etc.) as encrypted data.
[0038] The querying party is a fund marketing management platform, which initiates a query request for industry average data.
[0039] Coordinator: An independent third-party platform responsible for coordinating the computation process.
[0040] Throughout the data sharing process, technologies such as Multi-Party Secure Computation (MPC) and Anonymous Query can be used to ensure the privacy and security of data during transmission and computation. MPC allows multiple data providers to jointly compute data, such as calculating industry averages, without revealing their original plaintext information. All computations are performed in an encrypted state, ensuring that the original data remains local and does not leave the domain. The coordinating party also cannot access any party's detailed data. This embodiment can also employ technologies such as Oblivious Transfer (OT) to ensure that the querying party obtains the results of the joint computation without exposing its query ID (e.g., querying data from a specific channel).
[0041] In one embodiment, the step of acquiring industry panoramic data and performing benchmark calculations on the industry panoramic data to obtain industry average data further includes: Based on the aforementioned industry panoramic data, benchmark calculations are performed using statistical analysis and machine learning algorithms to obtain the industry average data.
[0042] Through the aforementioned secure sharing mechanism, this embodiment can acquire standardized, consistent data from multiple channels. The data dashboard calculates statistically significant industry averages, such as average UV, average purchase conversion rate, and average order value. These data serve as a benchmark for comparison with independent channel data, providing operators with a scientific horizontal comparison standard. By comparing their own data with industry benchmarks, operators can intuitively identify weaknesses in their marketing funnel. For example, if a channel's "click-through rate" is significantly lower than the industry average, it can be inferred that there are problems with user reach and content presentation on that channel.
[0043] Specifically, benchmark calculations can be performed on the collected industry-wide data using statistical analysis and machine learning algorithms. In terms of statistical analysis, descriptive statistics, such as calculating the mean, median, and mode, can be used to understand the central tendency of the data; variance and standard deviation can be used to measure the dispersion of the data; and correlation coefficients can be used to analyze the relationships between different data points. For example, analyzing the correlation between UV and purchase click-through rate in page tracking data can determine the connection between user visits and purchase intentions. Machine learning algorithms can further uncover potential patterns and regularities in the data. For instance, clustering algorithms can be used to segment customers, dividing them into different groups based on their transaction behavior and profile characteristics, and developing personalized marketing strategies for different groups; regression algorithms can be used to predict future industry trends, such as predicting the subscription amount and number of customers for gold ETF linked funds in the future, providing forward-looking guidance for fund companies' operational decisions. By combining statistical analysis with machine learning algorithms in benchmark calculations, more accurate and comprehensive industry average data can be obtained, providing a solid basis for data management and operational decisions of the marketing management platform.
[0044] In one embodiment, the step of performing intelligent analysis using a large language model based on the multi-source heterogeneous data and industry average data, and outputting corresponding data management strategies, includes: The multi-source heterogeneous data and / or industry average data are input into the large language model, and the large language model performs deep analysis and pattern recognition on the input data. Based on the analysis results and pattern recognition, targeted data management strategy suggestions are generated; The generated data management strategy recommendations will be output in the form of structured reports or visual charts.
[0045] This embodiment employs a large language model to analyze and process the input multi-source heterogeneous data and / or industry average data (such as marketing funnel data, customer profiles and transaction data labels, comparison results with industry benchmark data, real-time market information and investment research opinions, etc.). Leveraging its powerful natural language processing capabilities and deep learning algorithms, the large language model can automatically identify complex patterns, potential correlations, and anomalies in the data. For example, it can accurately identify significant changes in customer transaction behavior over different time periods or discover unique patterns in product preferences among specific customer groups. Based on these in-depth analysis and pattern recognition results, the large language model further generates highly targeted data management strategy recommendations. These recommendations not only cover conventional areas such as product recommendations and marketing campaign planning but also delve into refined management levels such as customer segmentation and risk warning. For instance, for a specific customer group, the model might suggest launching customized financial product portfolios or designing a series of interactive marketing activities aimed at increasing customer activity. Meanwhile, this embodiment presents strategy recommendations in the form of structured reports or visual charts. The structured reports, with their clear logical framework and well-organized chapters, elaborate on the background, objectives, specific content, and expected effects of the strategy recommendations. The visual charts, on the other hand, use intuitive graphics, colors, and animations to present complex data relationships and key strategy points in a concise and clear manner, greatly improving the efficiency and accuracy of information transmission.
[0046] In one embodiment, the data management method further includes: Historical data is used to train the large language model, and a supervised fine-tuning mechanism is used to fine-tune and optimize the trained and optimized large language model.
[0047] Before generating strategy recommendations using a large language model, it is thoroughly trained with historical data. Historical data can include unstructured data such as historical marketing reports, investment research content, customer service records, and market analysis documents. Using this unstructured data for model training allows the large language model to better understand the logical relationships behind the data. For example, by analyzing customer transaction data and page visit data from different channels over a period of time, the model can learn which factors have a significant impact on customer purchasing decisions and the differences in behavioral characteristics among different customer groups. After initial training, a supervised fine-tuning mechanism is used to further optimize the large language model. Supervised fine-tuning refers to targeted adjustments to the already trained model with labeled data. Through fine-tuning, the model can learn and understand professional knowledge in areas such as fund marketing, customer behavior, and market dynamics, thereby generating high-quality, highly accurate professional insights and recommendations.
[0048] This embodiment uses a Large Language Model (LLM) to generate in-depth, relevant, and actionable intelligent insights from the input structured data. Here, insight refers to interesting data patterns discovered from multidimensional data. The LLM in this embodiment can identify complex "compound insights," such as, "In a recent marketing campaign, while a specific channel had high unique visitors (UV), its purchase conversion rate was far below the industry average, and the main visitors to this channel were low-value customers." or "For a newly launched product, its average order value on channel A is much higher than on channel B. Data analysis shows that users on channel A are mainly reached through fund manager investor education videos, while users on channel B are mainly reached through market information pages." The LLM transforms these insights into specific, actionable marketing strategy recommendations. This is no longer a simple report, but an intelligent agent that proactively provides decision-making basis and optimization solutions. For example, based on the above insights, the LLM can provide the following recommendations: "We recommend optimizing the landing page content for Channel A by adding more investor education materials to improve conversion rates. We also recommend adjusting the advertising strategy to target high-value user groups more precisely." "For new products, we recommend strengthening the penetration of research and investment content on Channel B to increase users' willingness to buy and average order value."
[0049] This embodiment of the LLM application addresses the core problems of existing technologies by improving "marketing precision, marketing efficiency," and "optimization." It is not simply report analysis, but rather an intelligent decision support system capable of predicting potential customer needs, identifying potential customer value, and providing optimization strategies.
[0050] Figure 2 A schematic block diagram of a data management system 200 provided in an embodiment of the present invention, the system 200 including: The first acquisition unit 201 is used to acquire multi-dimensional data from different data sources and integrate the multi-dimensional data to obtain multi-source heterogeneous data. The second acquisition unit 202 is used to acquire industry panoramic data and perform benchmark calculations on the industry panoramic data to obtain industry average data. The intelligent analysis unit 203 is used to perform intelligent analysis based on the multi-source heterogeneous data and industry average data using a large language model, and output corresponding data management strategies.
[0051] In one embodiment, the first acquisition unit 201 includes: A multi-dimensional extraction unit is used to extract the multi-dimensional data from different data sources; wherein, the multi-dimensional data includes front-end page data, transaction system data, customer profile data, and external cooperation channel data; The data preprocessing unit is used to preprocess the multi-dimensional data to obtain the multi-source heterogeneous data.
[0052] In one embodiment, the data preprocessing unit includes: A standardization processing unit is used to standardize the multi-dimensional data using a Z-fraction scaling algorithm; The normalization processing unit is used to normalize the multi-dimensional data using the Min-Max normalization algorithm; An anomaly detection unit is used to perform outlier detection processing on the multi-dimensional data using a three-standard-deviation algorithm. The cleaning and desensitization unit is used to clean and desensitize the multi-dimensional data.
[0053] In one embodiment, the second acquisition unit 202 includes: The panoramic acquisition unit is used to acquire panoramic industry data from multiple related industry data providers by employing a multi-party secure computation model and an anonymous query algorithm.
[0054] In one embodiment, the second acquisition unit 202 further includes: The benchmark calculation unit is used to perform benchmark calculations based on the industry panoramic data, using statistical analysis and machine learning algorithms, to obtain the industry average data.
[0055] In one embodiment, the intelligent analysis unit 203 includes: The parsing and recognition unit is used to input the multi-source heterogeneous data and / or industry average data into the large language model, and to perform deep parsing and pattern recognition on the input data through the large language model; The strategy generation unit is used to generate targeted data management strategy suggestions based on the parsing results and pattern recognition. The strategy output unit is used to output the generated data management strategy recommendations in the form of structured reports or visual charts.
[0056] In one embodiment, the data management system 200 further includes: The training and optimization unit is used to train the large language model using historical data and to fine-tune and optimize the trained and optimized large language model using a supervised fine-tuning mechanism.
[0057] Since the embodiments of the apparatus and the embodiments of the method correspond to each other, please refer to the description of the embodiments of the method for the embodiments of the apparatus, which will not be repeated here.
[0058] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed, can perform the steps provided in the above embodiments. The storage medium may include various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0059] This invention also provides a computer device, which may include a memory and a processor. The memory stores a computer program, and when the processor calls the computer program in the memory, it can implement the steps provided in the above embodiments. Of course, the computer device may also include various network interfaces, power supplies, and other components.
[0060] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to in the method section. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
[0061] It should also be noted that, in this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
Claims
1. A data management method, characterized in that, include: Acquire multi-dimensional data from different data sources and integrate the multi-dimensional data to obtain multi-source heterogeneous data; Acquire industry panoramic data and perform benchmark calculations on the industry panoramic data to obtain industry average data; Based on the aforementioned multi-source heterogeneous data and industry average data, intelligent analysis is performed using a large language model, and corresponding data management strategies are output.
2. The data management method according to claim 1, characterized in that, The process of acquiring multi-dimensional data from different data sources and integrating and processing the multi-dimensional data to obtain multi-source heterogeneous data includes: The multi-dimensional data is extracted from different data sources; wherein, the multi-dimensional data includes front-end page data, transaction system data, customer profile data, and external cooperation channel data; The multi-dimensional data is preprocessed to obtain the multi-source heterogeneous data.
3. The data management method according to claim 2, characterized in that, The preprocessing of the multi-dimensional data to obtain the multi-source heterogeneous data includes: The multi-dimensional data is standardized using the Z-fraction scaling algorithm; The multi-dimensional data is normalized using the Min-Max normalization algorithm; The three-standard-deviation algorithm is used to perform outlier detection processing on the multi-dimensional data; The multi-dimensional data is then cleaned and de-identified.
4. The data management method according to claim 1, characterized in that, The process of acquiring industry panoramic data and performing benchmark calculations on the industry panoramic data to obtain industry average data includes: A multi-party secure computation model and anonymity query algorithm are used to obtain panoramic industry data from multiple related industry data providers.
5. The data management method according to claim 4, characterized in that, The process of acquiring industry panoramic data and performing benchmark calculations on the industry panoramic data to obtain industry average data also includes: Based on the aforementioned industry panoramic data, benchmark calculations are performed using statistical analysis and machine learning algorithms to obtain the industry average data.
6. The data management method according to claim 1, characterized in that, The method involves using a large language model for intelligent analysis based on the multi-source heterogeneous data and industry average data, and outputting corresponding data management strategies, including: The multi-source heterogeneous data and / or industry average data are input into the large language model, and the large language model performs deep analysis and pattern recognition on the input data. Based on the analysis results and pattern recognition, targeted data management strategy suggestions are generated; The generated data management strategy recommendations will be output in the form of structured reports or visual charts.
7. The data management method according to claim 1, characterized in that, Also includes: Historical data is used to train the large language model, and a supervised fine-tuning mechanism is used to fine-tune and optimize the trained and optimized large language model.
8. A data management system, characterized in that, include: The first acquisition unit is used to acquire multi-dimensional data from different data sources and integrate the multi-dimensional data to obtain multi-source heterogeneous data. The second acquisition unit is used to acquire industry panoramic data and perform benchmark calculations on the industry panoramic data to obtain industry average data. The intelligent analysis unit is used to perform intelligent analysis based on the multi-source heterogeneous data and industry average data, using a large language model, and output corresponding data management strategies.
9. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the data management method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the data management method as described in any one of claims 1 to 7.