A product development method based on public data reorganization and innovation scene mining
Patent Information
- Application Number
- CN202610849069.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-08-21
AI Technical Summary
然而,目前这些数据多以原始、分散的形式存在,实际利用率不足15%
[0020]本发明提供了一种基于公共数据重组与创新场景挖掘的产品开发方法。具备以下有益效果:
Smart Images

Figure CN122616413A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data processing and product innovation design technology, specifically a product development method based on public data restructuring and innovation scenario mining. Background Technology
[0002] With the advancement of digital technology, a vast amount of public data resources have been made available, covering multiple fields such as traffic flow data, environmental quality monitoring data, population statistics, economic indicators, and geographic information data. Statistics show that by 2025, my country had over 200 government data open platforms at the prefecture-level city level and above, with over 300,000 open datasets. However, currently, this data largely exists in raw and scattered forms, with an actual utilization rate of less than 15%.
[0003] In existing technologies, product development mainly relies on internal enterprise data, market research, or expert experience, which has the following technical shortcomings: First, there is a lack of systematic mechanisms for utilizing cross-domain data. Traditional product innovation methods (such as brainstorming and the Delphi method) rely heavily on subjective human experience and cannot systematically mine potential cross-domain relationships from massive amounts of public data. For example, whether there is a causal relationship between traffic congestion data and air quality data, or whether there is spatial coupling between population migration data and the distribution of business outlets, these implicit relationships are difficult to discover manually.
[0004] Second, data-driven product development lacks a scenario-oriented approach. Existing data analysis methods (such as data visualization and report statistics) can only present the data itself and cannot automatically transform data characteristics into product feature definitions with commercial value. There is a clear logical gap between "data insight" and "product definition".
[0005] Third, the evaluation of innovative scenarios lacks automated mechanisms. Even when certain data correlations are discovered, it is difficult to quickly determine whether a commercially viable solution already exists, leading to duplication of inventions and waste of resources.
[0006] In summary, the existing technology lacks a method that can systematically utilize public data, automatically mine innovative scenarios, and directly generate product development plans. This is the technical problem that this invention aims to solve. Summary of the Invention
[0007] (a) Technical problems to be solved
[0008] To address the shortcomings of existing technologies, this invention provides a product development method based on public data restructuring and innovative scenario mining, which solves the following problems:
[0009] (II) Technical Solution
[0010] To achieve the above objectives, the present invention provides the following technical solution: A product development method based on public data reorganization and innovative scenario mining includes an integrated monitoring and perception layer for real-time monitoring of urban weather, rainfall, water level, flow, water depth and underground pipe network status, and to obtain multi-source monitoring data of urban flood influencing factors; The digital twin data base layer, connected to the integrated monitoring and sensing layer, is used to aggregate, manage, fuse, and store multi-source heterogeneous data to build a standardized flood-themed database. The digital twin foundation layer, built upon the digital twin data base layer, includes a CIM platform, a model platform, and a knowledge platform, used to construct an integrated digital twin scenario of the city's above-ground and underground areas; A flood simulation model engine, deployed in the model platform, includes a hydrological and hydrodynamic coupling model and a machine learning proxy model, used for real-time simulation of the flood evolution process based on real-time monitoring data and meteorological forecast data; The four-prevention function application layer is connected to the flood simulation model engine and the digital twin base layer, and is used to realize the functions of forecasting, early warning, simulation and contingency planning; The visualization and interaction module is connected to the digital twin base layer and the four-pre-function application layer, and is used to visualize the flood situation and the simulation results in a three-dimensional manner in the digital twin scenario.
[0011] As a further preferred embodiment of the present invention, the flood simulation model engine adopts a dual-model collaborative computing architecture, including: First mode: High-precision offline mode, based on the hydrological and hydrodynamic coupling model to carry out refined flood simulation, with a calculation grid accuracy of ≤10m×10m, and the simulation results are used as training label data for machine learning proxy models; Second mode: Real-time online mode, which performs real-time inference based on the trained machine learning agent model, with an inference update cycle of ≤5min; The workflow of the dual-mode collaborative computing architecture is as follows: daily model calculations are performed using high-precision offline mode based on the latest data base, generating a simulation database for incremental training of the machine learning agent model; during the real-time monitoring phase, the machine learning agent model is invoked for rapid inference; when extreme rainstorms are forecast, high-precision offline mode is automatically triggered for refined verification calculations.
[0012] As a further preferred embodiment of the present invention, the CIM platform includes: The terrain modeling module uses airborne LiDAR point cloud data to generate a high-precision digital elevation model (DEM), filters the point cloud data, classifies ground points and non-ground points, and constructs an irregular triangular network to form a terrain model. The underground pipe network modeling module integrates pipeline data and node data to construct a three-dimensional topology model of the underground drainage pipe network. Pipeline data includes starting point coordinates and elevation, pipe length, pipe type, and cross-sectional area attributes. Node data includes node coordinates, node elevation, and appurtenance type attributes. The 3D rendering engine module adopts a cloud rendering architecture to distribute 3D rendering tasks to cloud servers for processing.
[0013] As a further preferred embodiment of the present invention, the knowledge platform includes: The historical flood case database includes information on rainfall events, inundation extent, damage, and response measures for historical flood disasters. A flood control rules knowledge graph is constructed, which includes rainfall thresholds, early warning rules, dispatching rules, and emergency plans. The similar rainfall pattern matching engine recommends similar rainfall events and corresponding response experiences from a historical case database based on the current rainfall forecast and similarity analysis.
[0014] As a further preferred embodiment of the present invention, in the four pre-functional application layer: The forecast module is used to predict the evolution of urban flooding in the next 1-48 hours, including forecasts of inundation range, water depth, and receding time. The early warning module is used to automatically determine the four levels of early warning (red, orange, yellow, and blue) based on the forecast results, and to identify the uncertain population within the scope of the early warning based on the population heat map in the digital twin scenario and push the early warning accordingly. The simulation module is used to dynamically simulate and rehearse the evolution of floods and emergency response plans in a digital twin scenario, and supports comparison of multiple plans; The contingency plan module is used to automatically generate flood control emergency plans by integrating the results of pre-drill analysis, historical case database, and flood control rule knowledge graph.
[0015] As a further preferred embodiment of the present invention, it also includes a terrain dynamic update module, which is used to calculate the erosion assessment value and sedimentation assessment value of each local area based on real-time monitored water level and flow velocity data, dynamically update the terrain elevation data in the digital twin model, and feed the updated terrain data back to the hydrological and hydrodynamic coupling model in real time. The equipment reliability assessment module is used to construct an equipment compressive performance assessment formula based on real-time monitoring data, calculate a comprehensive assessment index in combination with environmental factors, and generate an equipment anomaly warning when the assessment index is lower than the threshold.
[0016] As a further preferred embodiment of the present invention, the integrated monitoring and sensing layer includes a meteorological monitoring unit, a hydrological monitoring unit, a water accumulation monitoring unit, and a pipeline monitoring unit; the water accumulation monitoring unit deploys electronic water gauges or buried water level sensors in low-lying road sections and underpasses; the pipeline monitoring unit deploys flow meters and level gauges at key nodes of the drainage pipeline network.
[0017] As a further preferred embodiment of the present invention, the machine learning proxy model adopts an LSTM-CNN hybrid architecture or a Transformer architecture, with the input layer receiving the rainfall sequence and initial conditions, and the output layer predicting the water depth change process at each grid point.
[0018] As a further preferred embodiment of the present invention, the method of operating the system includes the following steps: S1 collects multi-source flood monitoring data in real time through an integrated monitoring and sensing layer; S2 aggregates the collected data to the digital twin data base layer for data governance and fusion; S3, based on the digital twin foundation layer, constructs an integrated digital twin scenario for urban above-ground and underground environments; S4 invokes the flood simulation model engine to perform real-time simulation of flood evolution based on real-time monitoring data and meteorological forecast data; S5, based on the simulation results, performs forecast analysis, early warning issuance, scheme simulation and contingency plan generation through the four-prevention function application layer; S6 uses a visualization and interactive module to display the flood situation and projection results in a 3D digital twin scenario.
[0019] (III) Beneficial Effects
[0020] This invention provides a product development method based on public data restructuring and innovative scenario mining. It has the following beneficial effects: Breaking down data silos and creating incremental data value. By mining cross-domain association rules, previously unrelated traffic data, environmental data, population data, and economic data are systematically reorganized, creating new data dimensions that transcend the value of a single data source. Compared to manual analysis, this invention improves cross-domain association discovery efficiency by approximately 20 times and reduces the false negative rate to below 15%.
[0021] Lowering the barrier to innovation and automating product definition: This invention establishes an automated mapping mechanism from "data characteristics" to "functional definitions," transforming vague data insights into clear product requirements. Developers do not need cross-domain expertise to obtain data-supported product directions, and the cost of R&D trial and error is expected to be reduced by more than 60%.
[0022] Dynamic adaptability and continuous product solution iteration. As public data is continuously updated and its accessibility expands, the method of this invention can automatically trigger re-mining, ensuring that product solutions remain synchronized with the latest social operating rules. Simultaneously, the feedback mechanism enables the system to learn from evaluation results, continuously improving the accuracy of scenario selection.
[0023] Avoid redundant inventions and focus on genuine innovation. By automatically comparing existing patents and commercial products through a scene boundary recognition model, it effectively avoids developing existing technical solutions and focuses innovation resources on truly untapped scenarios. Attached Figure Description
[0024] Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 This is a schematic diagram of the system framework principle of the present invention. Detailed Implementation
[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0026] Please see Figure 1-2 This invention provides a technical solution: a product development method based on public data restructuring and innovative scenario mining, comprising the following steps: Step S1: Standardized collection of multi-source public data and construction of knowledge graph Establish application programming interfaces (APIs) with at least one public data open platform to collect structured and unstructured data. These public data open platforms include, but are not limited to: local government data open platforms (such as the Shanghai Municipal Government Data Service Network), meteorological data platforms (such as the China Meteorological Data Network), transportation data platforms (such as local transportation operation monitoring centers), and environmental monitoring data platforms (such as online air quality monitoring platforms).
[0027] The collected data formats include: structured data in JSON format, tabular data in CSV format, and semi-structured data in XML format.
[0028] The collected data was processed as follows: (1) Cleaning process: Remove obvious outliers, such as records with a traffic congestion index greater than 10 (normal range is 0-10) and records with negative PM2.5 concentration; use linear interpolation to fill missing values, and remove data segments with more than 3 consecutive missing time points.
[0029] (2) Entity alignment processing: unify the different representations of the same entity in different data sources. For example, “Zhongshan East Road” in traffic data and “Zhongshan Road East Section” in environmental data need to be aligned to the same geographic entity. Entity alignment adopts a string matching algorithm based on edit distance, and the similarity threshold is set to 0.85.
[0030] (3) Normalization: Unify data of different dimensions to the same numerical range. The Min-Max normalization method is adopted, and the formula is: X_norm=(X-X_min) / (X_max-X_min), which maps all numerical features to the interval [0,1].
[0031] A public data knowledge graph is constructed, incorporating time, space, and topic dimensions. The knowledge graph is stored in triplet (entity-relationship-entity) format, and its data structure is shown in the table below:
[0032] Entity types include: geographic entities (road segments, monitoring stations, administrative regions), time entities (time periods, dates), and numerical entities (congestion index, PM2.5 concentration levels). Relationship types include: temporal relationships ("precedes" indicates earlier in time, "follows" indicates later in time), spatial relationships ("located_in" indicates located within, "within_distance" indicates within a distance), and statistical relationships ("correlates_with" indicates a statistical correlation).
[0033] Step S2: Cross-domain data reorganization based on multidimensional association rule mining An improved FP-Growth (Frequent Pattern Growth) association rule algorithm is used to mine cross-domain implicit associations between different entities in the knowledge graph. Compared with the traditional Apriori algorithm, the FP-Growth algorithm has higher computational efficiency, requiring only two database scans and no candidate set generation, making it particularly suitable for large-scale knowledge graph data.
[0034] The evaluation metrics and calculation formulas for association rules are as follows: Support: Support(A→B)=P(A∪B)=count(A∪B) / N, where N is the total number of transactions, representing the frequency of the rule in the dataset.
[0035] Confidence: Confidence(A→B)=P(B|A)=count(A∪B) / count(A), representing the probability of B occurring given that A has occurred.
[0036] Lift: Lift(A→B) = Confidence(A→B) / Support(B). A lift greater than 1 indicates a positive correlation between A and B, equal to 1 indicates they are independent, and less than 1 indicates a negative correlation.
[0037] The association rule filtering thresholds set in this invention are: minimum support ≥ 0.05, minimum confidence ≥ 0.6, and minimum lift ≥ 1.2. Only association rules that simultaneously meet all three conditions are retained.
[0038] Semantic recombination is performed on two or more data entities from different primary business domains that meet the conditions to generate a "composite data feature set". The primary business domains include at least two of the following: transportation, environment, population, economy, and geographic information.
[0039] The specific operations of semantic reorganization include combining the antecedent (A) and consequent (B) from the mined association rules to form a new composite feature vector. The dimension of this composite feature vector is the sum of the dimensions of the two original feature vectors, and the evaluation index of the association rule is added as a weight.
[0040] Step S3: Innovative Scene Mining and Boundary Recognition Based on Density Clustering First, the composite data feature set is mapped to a predefined social demand labeling system. This social demand labeling system is a pre-constructed multi-level classification system, the structure of which is shown in the table below: First-level label, second-level label, label coding explanation Intelligent Traffic Congestion Management T-01: Addressing the Needs Related to Alleviating Traffic Congestion Smart Transportation Parking Guidance T-02: Parking Space Search and Reservation Related Needs Environmental governance pollution source tracing E-01: Related needs for tracing pollutant sources Environmental Governance Dynamic Early Warning E-02: Related Needs for Early Warning of Pollution Incidents Public Safety Emergency Response S-01: Requirements for Emergency Response to Sudden Incidents Commercial Site Optimization Recommendation B-01: Related Needs for Commercial Network Layout Optimization Public service resource allocation P-01 Optimization of resource allocation in areas such as healthcare and education The mapping method employs text semantic similarity calculation based on BERT (Bidirectional Encoder Representations from Transformers). The data description text in the composite data feature set is vectorized with the description text of each label, cosine similarity is calculated, and the label with the highest similarity is selected as the annotation result for that feature set.
[0041] Secondly, the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) density clustering algorithm is used to cluster the mapped feature data, generating candidate application scenario clusters. The main parameter settings of the DBSCAN algorithm are as follows: Neighborhood radius ε=0.3: This means that two samples within a distance of 0.3 are considered to be adjacent.
[0042] Minimum neighborhood number MinPts=5: This means that a core point must contain at least 5 samples in its neighborhood.
[0043] The core idea of the DBSCAN algorithm is: starting from any unvisited sample point, find all sample points within its ε-neighborhood. If the number of samples in this neighborhood is greater than or equal to MinPts, a cluster is formed, and the cluster is recursively expanded; otherwise, the point is marked as noise. This algorithm can automatically identify clusters of arbitrary shapes without needing to pre-specify the number of clusters, making it particularly suitable for scene mining where the cluster shape is uncertain and the number of clusters is unknown.
[0044] Next, scenarios with existing commercial solutions are excluded using a scene boundary recognition model, thus identifying "vacuum scenarios" with innovative value. The execution flow of the scene boundary recognition model is as follows: (1) Perform patent database search and comparison on candidate application scenarios. Construct a search expression and match the keywords of the candidate scenarios (extracted from the composite data feature set) with the titles, abstracts and claims in the patent database. The search results return the top 100 most relevant patents.
[0045] (2) Conduct a search and comparison of candidate application scenarios using a commercial product database. The commercial product database includes: mainstream app stores (APP Store, Google Play, Huawei App Market), SaaS product platforms (such as SaaS Review Network), and enterprise information query platforms (such as Tianyancha, Qichacha). The search method is similar to that of patent search.
[0046] (3) Overlap Calculation. The overlap between the candidate scenario and the existing solution's technical path is calculated using the Jaccard similarity coefficient. The formula is: J(A,B)=|A∩B| / |A∪B|, where A is the set of technical features of the candidate scenario (extracted from the composite data feature set), and B is the set of technical features of the existing solution (extracted from patent documents or product descriptions). Technical features include dimensions such as data source type, algorithm type, output format, and application field.
[0047] (4) Threshold judgment. When J(A,B)<0.3 (i.e., the overlap is less than 30%), the candidate scenario is marked as a "vacuum scenario", indicating that the scenario has high innovation value; when J(A,B)≥0.7, it is marked as a "saturated scenario" and it is recommended to abandon it; when 0.3≤J(A,B)<0.7, it is marked as a "competitive scenario" and it is recommended to further analyze the differentiation space.
[0048] Step S4: Generate a product prototype solution based on function mapping First, based on the vacuum scenario and the corresponding composite data feature set, a preset "data feature-technical solution" mapping library is invoked to perform functional mapping and generate a product functional requirement list.
[0049] The "data feature-technical solution" mapping library is a pre-built relational database, and its data structure is shown in the table below:
[0050] The specific process of function mapping is as follows: each feature in the composite data feature set is matched with the feature_label in the mapping library. If the match is successful, the tech_component_id corresponding to the feature is added to the function requirement list, and its calling interface specification is recorded.
[0051] Secondly, based on the aforementioned list of functional requirements, a pre-defined modular technology component library is invoked to automatically generate a product development prototype. This modular technology component library encapsulates various mature technology components, including but not limited to: Data processing components: data cleaning component, missing value imputation component, outlier detection component; Feature engineering components: feature extraction component, feature selection component, feature dimensionality reduction component; Model components: LSTM time series prediction model, random forest classification model, DBSCAN clustering model; Visualization components: map visualization components, chart visualization components, dashboard components; Deployment components: API encapsulation components, front-end UI components, push service components; The product development prototype is generated using a template-based approach. The preset product prototype template includes the following fields: Product Name (automatically generated based on scenario tags), Product Positioning Description (automatically generated based on composite data feature sets), Core Functional Modules (automatically generated based on the functional requirement list), Technical Architecture Diagram (automatically generated based on called components), Data Flow Diagram (automatically generated based on data flow direction), and Development Workload Estimate (automatically estimated based on the number of components).
[0052] Step S5: Product Value Assessment and Iterative Output A multi-dimensional evaluation model is constructed to evaluate the product development prototype solution. The multi-dimensional evaluation model includes four dimensions, and the evaluation criteria and weights for each dimension are as follows: The overall score is calculated using a weighted summation method, with the formula: Score_total = Score_tech × 0.25 + Score_data × 0.20 + Score_business × 0.35 + Score_social × 0.20.
[0053] When the overall score is ≥70, the product solution is marked as "Recommended for Development" and output to the development terminal; when the overall score is <70, it is marked as "Needs Optimization" and the system automatically generates optimization suggestions (such as supplementing data sources, replacing technical components, etc.).
[0054] Simultaneously, the boundary recognition threshold in step S3 and the technical component library in step S4 are updated through the feedback interface. The specific update mechanism is as follows: for product solutions that pass the evaluation, their corresponding scenario features are recorded as "verified scenarios" to reduce the boundary recognition threshold for subsequent similar scenarios (from 30% to 20%); for solutions that fail the evaluation, the reasons for failure are analyzed, and if the technical component is unavailable, the component is marked as "to be developed" and a research and development task is triggered.
[0055] The data platform module is used to perform standardized collection of multi-source public data and knowledge graph construction. The data platform module includes an API gateway unit, a data cleaning unit, an entity alignment unit, and a normalization processing unit. The API gateway unit is configured with multiple data source connectors, each corresponding to a data interface of a public data open platform, supporting multiple protocols such as RESTful API, SOAP API, and OData. The output of the API gateway unit is connected to the input of the data cleaning unit. The data cleaning unit has a built-in outlier detector and a linear interpolation filler to remove data records that are significantly outside the normal range and fill in missing values. The output of the data cleaning unit is connected to the input of the entity alignment unit, which has a built-in string matching algorithm module based on edit distance to unify different representations of the same entity from different data sources into a standardized entity name. The output of the entity alignment unit is connected to the input of the normalization processing unit, which uses the Min-Max normalization method to unify data of different dimensions into the [0,1] interval. The output of the data platform module is connected to a knowledge graph database, which is implemented using a graph database (such as Neo4j) and is used to store public data knowledge graphs containing time, space and topic dimensions. The knowledge graph is stored in the form of triples (entity-relationship-entity).
[0056] The analysis and reorganization module is used to perform cross-domain data reorganization. This module includes an FP-Growth algorithm engine unit and a semantic reorganization unit. The input of the FP-Growth algorithm engine unit is connected to the output of the knowledge graph database. Internally, it encapsulates an FP-Tree builder and a frequent pattern miner to discover implicit cross-domain relationships between different entities. The output of the FP-Growth algorithm engine unit is connected to the input of the semantic reorganization unit. The semantic reorganization unit has a built-in feature vector combiner used to semantically reorganize two or more data entities from different primary business domains that meet the following criteria: support ≥ 0.05, confidence ≥ 0.6, and lift ≥ 1.2, generating a composite data feature set.
[0057] The scenario incubation module is used to perform innovative scenario mining and boundary recognition. The scenario incubation module includes a BERT semantic mapper unit, a DBSCAN clustering unit, and a boundary recognition unit. The input of the BERT semantic mapper unit is connected to the output of the semantic recombination unit. Internally, it loads a pre-trained "bert-base-chinese" model and a predefined social demand label system database. This is used to vectorize the data description text in the composite data feature set and the description text of each label, calculate the cosine similarity, and select the label with the highest similarity as the labeling result. The output of the BERT semantic mapper unit is connected to the input of the DBSCAN clustering unit. The DBSCAN clustering unit internally encapsulates a density clustering algorithm module, set with a neighborhood radius ε=0.3 and a minimum number of neighborhood points MinPts=5, used to cluster the mapped feature data to generate candidate application scenario clusters. The output of the DBSCAN clustering unit is connected to the input of the boundary recognition unit. The boundary identification unit includes a patent retrieval subunit, a commercial product retrieval subunit, and a Jaccard similarity calculation subunit. It is used to perform patent database retrieval and comparison and commercial product database retrieval and comparison on candidate application scenarios, calculate the technical path overlap, and mark the candidate scenario as a vacuum scenario when the overlap is less than 30%.
[0058] The design output module is used to generate a product prototype solution. The design output module includes a function mapper unit, a component library interface unit, and a template filling unit. The input of the function mapper unit is connected to the output of the boundary identification unit. Internally, it stores a "data feature-technical solution" mapping library (this mapping library is a pre-built relational database; each record contains a data feature tag field, a corresponding mature technology component identifier field, a technology component call interface specification field, and a confidence score field for the mapping relationship). This library is used to perform function mapping based on the vacuum scenario and generate a product functional requirement list. The output of the function mapper unit is connected to the input of the component library interface unit. The component library interface unit is used to call corresponding components in the modular technology component library via a RESTful API. Its output is connected to the input of the template filling unit. The template filling unit internally stores a product prototype template, which includes a product name field, a product positioning description field, a core functional module list field, a technical architecture diagram placeholder, a data flow diagram placeholder, and a development workload estimation field. This template is used to fill the functional requirement list and technical component call information into the product prototype template, automatically generating a product development prototype solution.
[0059] The evaluation feedback module is used to perform product value evaluation and iterative output. The evaluation feedback module includes a multi-dimensional evaluator unit and a feedback interface unit. The input of the multi-dimensional evaluator unit is connected to the output of the template filling unit, and it internally stores a four-dimensional evaluation weight configuration table (technical feasibility 25%, data real-time performance 20%, commercial value 35%, social benefits 20%), used to weight and score the product development prototype scheme. The output of the multi-dimensional evaluator unit is connected to the development terminal to output a product scheme with a comprehensive score ≥70 points, and also to the input of the feedback interface unit. The output of the feedback interface unit is connected to the threshold adjustment terminal of the boundary recognition unit and the mapping library update terminal of the function mapper unit, respectively. It is used to feed back the scene features corresponding to the evaluated and qualified schemes to the boundary recognition unit to reduce the boundary recognition threshold for similar scenes (from 30% to 20%), and to feed back the mapping relationship between successfully invoked technical components and data feature labels to the function mapper unit to update the confidence of the mapping library (increasing it by 0.01-0.03).
[0060] Computer-readable storage media can be implemented using any one or more of the following types of physical media, depending on the deployment scenario:
[0061] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
[0062] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
[0063] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A product development method based on public data restructuring and innovative scenario mining, characterized in that, Includes the following steps: Step S1: Standardized collection and knowledge graph construction of multi-source public data: Establish an interface with at least one public data open platform, collect multi-source heterogeneous data, clean, align and normalize the collected data, and construct a public data knowledge graph that includes time, space and topic dimensions. Step S2: Cross-domain data reorganization based on multi-dimensional association rule mining: The improved FP-Growth association rule algorithm is used to mine cross-domain implicit associations between different entities in the knowledge graph. Two or more data entities from different first-level business domains with an association degree higher than a preset threshold are semantically reorganized to generate a composite data feature set. Step S3: Innovative scenario mining and boundary identification based on density clustering: The composite data feature set is mapped to a predefined social demand label system, candidate application scenario clusters are generated using the DBSCAN density clustering algorithm, and scenarios with existing commercial solutions are excluded by the scenario boundary identification model, and vacuum scenarios with innovative value are selected. Step S4: Product prototype solution generation based on function mapping: Based on the vacuum scenario, combined with the corresponding composite data feature set, function mapping is performed using the preset "data feature-technical solution" mapping library to generate a product function requirement list, and the modular technology component library is called to automatically generate a product development prototype solution. Step S5: Product Value Assessment and Iterative Output: Construct a multi-dimensional assessment model to evaluate the product development prototype solution, output the qualified product solution to the development terminal, and update the boundary identification threshold of step S3 and the technical component library of step S4 through the feedback interface.
2. The product development method based on public data restructuring and innovation scenario mining according to claim 1, characterized in that: The cross-domain association rule mining in step S2 specifically includes: identifying temporal causal associations or spatial inclusion associations between at least two data entities from different primary business domains, wherein the primary business domains include at least two from the transportation domain, environment domain, population domain, economic domain, and geographic information domain.
3. The product development method based on public data restructuring and innovative scenario mining according to claim 1, characterized in that: The scene boundary recognition model in step S3 includes: performing patent database search and comparison and commercial product database search and comparison on candidate application scenarios. When the overlap between the candidate scenario and the technical path of the existing solution is less than a preset threshold of 30%, it is marked as a vacuum scenario.
4. The product development method based on public data restructuring and innovative scenario mining according to claim 1, characterized in that: The "data feature-technical solution" mapping library in step S4 is a pre-built association database, in which each record contains a data feature label, the corresponding mature technology component identifier, and the calling interface specification of the technology component.
5. A product development method based on public data restructuring and innovative scenario mining according to claim 1, characterized in that: The multi-dimensional evaluation model in step S5 includes: technical feasibility dimension, data real-time dimension, business value dimension, and social benefit dimension. Each dimension uses a weighted scoring method to calculate the comprehensive score.
6. A product development system based on public data restructuring and innovative scenario mining, characterized in that, include: The data platform module is used to execute step S1 in claim 1; The analysis and recombination module is used to perform step S2 in claim 1; The scenario incubation module is used to execute step S3 in claim 1; Design an output module to perform step S4 in claim 1; An evaluation feedback module is used to perform step S5 of claim 1.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 5.