Multi-source heterogeneous data driven artificial intelligence integrated software development system oriented to small and medium-sized enterprises

By designing a multi-source heterogeneous data-driven artificial intelligence integrated software development system for small and medium-sized enterprises, using CNN and genetic algorithms for adaptive learning and optimization, the problem of existing systems lacking update and adaptability is solved, and efficient data processing and model prediction performance is achieved.

CN119987756AInactive Publication Date: 2025-05-13HANGZHOU HUASHI ENTERPRISE MANAGEMENT CONSULTING CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510069864.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing software development systems lack effective update and adaptive learning mechanisms, cannot conduct online learning based on real-time data feedback, have weak ability to adapt to environmental changes, and feature extraction relies on manual selection, which is easy to introduce subjective bias and is inefficient.

Method used

Design a development system for multi-source heterogeneous data-driven artificial intelligence integrated software for small and medium-sized enterprises, including data crawling, multi-source data access, data preprocessing, model training, model evaluation, software integration and deployment, user interface and interaction, system management and maintenance and other modules. Adaptive learning is used for CNN algorithm and optimization is combined with genetic algorithms to achieve online update of the model and environmental adaptation.

Benefits of technology

By cleaning noise and outliers, improve data quality, use CNN to automatically extract features, avoid subjective bias, and improve the learning ability and prediction performance of the model. Multi-objective optimization and adaptive learning of the model are realized, and prediction accuracy and data processing efficiency are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987756A_ABST
    Figure CN119987756A_ABST
Patent Text Reader

Abstract

The invention discloses a development system of multi-source heterogeneous data driven artificial intelligence integrated software for medium and small enterprises. According to the method, the genetic algorithm is combined for optimization, the output of the deep learning model is used as the initial population, and the fitness function comprehensively considering the effect and the resource consumption is defined, so that multi-objective optimization is realized, and a parameter combination better meeting the actual demand is found. A model updating and self-adaptive learning mechanism ensures that the model can perform online learning according to data fed back in real time, environment changes are adapted in time, and the prediction accuracy is continuously improved. Therefore, the integration of the model training module not only improves the efficiency and quality of data processing, but also enhances the learning ability and prediction performance of the model, achieves the multi-objective optimization and adaptive learning, provides more intelligent and precise decision support and service for small and medium-sized enterprises, and improves the user experience. Innovation and development of small and medium-sized enterprises in the digital transformation process are effectively promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of software development for small and medium-sized enterprises, and specifically is a development system for multi-source heterogeneous data-driven artificial intelligence integrated software for small and medium-sized enterprises. Background Art

[0002] SME integration software is a comprehensive management software tailored for SMEs. It integrates various business processes of enterprises such as procurement, sales, inventory, finance, human resources and other modules to achieve data sharing and business collaboration, helping enterprises improve management efficiency, reduce operating costs, and optimize resource allocation. This kind of software usually has the characteristics of strong ease of use, fast deployment, and flexible expansion, which can meet the needs of SMEs at different stages of development. Integration software reduces human errors and improves data processing speed through standardized processes and automated operations. At the same time, it provides real-time and accurate reports and data analysis to provide strong support for enterprise decision-making. With the rapid development of information technology, big data and artificial intelligence technology have gradually become the key driving force for enterprise innovation and development. Especially in small and medium-sized enterprises, how to effectively use multi-source heterogeneous data and improve business efficiency and decision-making capabilities through artificial intelligence technology has become an urgent problem to be solved.

[0003] However, the existing software development system lacks effective update and adaptive learning mechanisms, cannot perform online learning based on real-time data feedback, and has a weak ability to adapt to environmental changes. At the same time, feature extraction often relies on manual selection, which is prone to subjective bias, has low efficiency, and cannot guarantee the quality of feature vectors. Summary of the invention

[0004] The purpose of the present invention is to provide a development system for multi-source heterogeneous data-driven artificial intelligence integrated software for small and medium-sized enterprises in order to solve the above-mentioned problems.

[0005] The technical solution adopted by the present invention is as follows: a development system for multi-source heterogeneous data-driven artificial intelligence integrated software for small and medium-sized enterprises, the system comprising: a data crawling submodule, a multi-source data access submodule, a data preprocessing module, a model training module, a model evaluation module, a software integration and deployment module, a user interface and interaction module, and a system management and maintenance module;

[0006] The software integration and deployment module is internally provided with a software framework integration submodule, an API generation and management submodule, and a deployment and monitoring submodule;

[0007] The data crawling submodule and the multi-source data access submodule serve as data entry points to collect data from different sources such as the Internet, databases, APIs, file systems, and IoT devices, and transmit the data to the data preprocessing module;

[0008] The data preprocessing module cleans, converts and performs feature engineering on the data to ensure the quality and consistency of the data and prepare for subsequent model training. The processed data enters the model training module.

[0009] The model training module uses integrated machine learning and deep learning algorithms to train the model, and optimizes the model performance through the hyperparameter tuning submodule; the trained model will be moved to the model evaluation module.

[0010] The model evaluation module performs performance evaluation, visualization analysis, and model interpretation to ensure the accuracy and interpretability of the model;

[0011] The software integration and deployment module integrates the model with the software framework, generates an API interface, supports one-click deployment to the cloud platform or local server, and provides real-time monitoring capabilities;

[0012] The user interface and interaction module provides a visual interface, user management and task scheduling functions, allowing users to interact with the system conveniently;

[0013] The system management and maintenance module is responsible for system configuration, log management, security and backup to ensure stable operation of the system and data security.

[0014] In a preferred embodiment, the data crawling submodule first defines the crawling target, including a list of related websites and product information to be crawled; then, the developer configures the crawler task in the data crawling submodule, sets the crawling frequency, data storage path, user agent and IP pool to avoid being blocked by the target website; after the configuration is completed, the crawler automatically starts running, accesses the target website according to the preset rules, parses the web page content, extracts the product information, and stores it in the specified location; the data analyst views the crawling results in real time to ensure the accuracy and completeness of the data; once data anomalies or crawling failures are found, the system will automatically send an alarm to notify relevant personnel so that the crawler strategy can be adjusted or the problem can be fixed in time;

[0015] The multi-source data access submodule configures data source links in the multi-source data access submodule, including connections to relational databases, real-time data streams accessed through API interfaces, and regularly uploaded CSV files; during the configuration process, the frequency of data extraction, data conversion rules, and target storage for data loading are defined; after the configuration is completed, the module automatically executes the data extraction, conversion, and loading processes to integrate data from different sources into a unified data warehouse; business analysts use these integrated data for analysis to generate various reports and insights; if there are delays or errors in the data access process, the system management and maintenance module will record relevant logs and notify data engineers to conduct investigations and repairs.

[0016] In a preferred embodiment, the data preprocessing module uses a minimum-maximum standardization method to perform data standardization processing, and the specific steps are as follows:

[0017] S1. Calculate the maximum and minimum values: For each feature, find the maximum and minimum values ​​in the Internet security information dataset;

[0018] S2. Apply transformation formula: For each feature value x in the Internet security information dataset, use the following formula for transformation:

[0019]

[0020] Among them, x norm is the transformed value;

[0021] S3. Transform the Internet security information dataset: Apply the above transformation to each feature of the entire dataset; after minimum-maximum normalization, all features will be scaled to the range of 0 to 1.

[0022] In a preferred embodiment, the model training module uses a CNN algorithm to perform adaptive learning based on historical data and real-time feedback from small and medium-sized enterprises, specifically including the following steps:

[0023] Data preprocessing: clean the collected SME data to remove noise and outliers; normalize the SME data and scale all features to the same scale;

[0024] Feature extraction: Use deep learning networks to automatically extract features of SME data; output feature vectors,

[0025] Used for subsequent optimization processing;

[0026] Deep learning model training: Design a deep learning model with the extracted feature vector as input and the recommended parameters of SME data as output; use historical SME data to train the model and optimize the model parameters through the back propagation algorithm;

[0027] Genetic algorithm optimization: Use the output parameters of the deep learning model as the initial population of the genetic algorithm; define the fitness function to reflect the effect and resource consumption of SME data; through the iterative search of the genetic algorithm,

[0028] Find the optimal parameter combination;

[0029] Model update and adaptive learning: Use online learning strategies to update deep learning models based on real-time feedback data, enabling the model to adapt to environmental changes and improve prediction accuracy;

[0030] The data normalization formula is:

[0031] Among them, x is the original data, x′ is the normalized data, μ is the mean of the data,

[0032] σ is the standard deviation of the data;

[0033] · The convolutional layer formula of CNN is:

[0034] · The convolutional layer formula is:

[0035]

[0036] Among them, I is the input SME data, K is the convolution kernel, and (i, j) is the position of the output SME data;

[0037] The pooling layer formula for maximum pooling is:

[0038]

[0039] Among them, P is the data feature of SMEs after pooling, and R is the data feature set of SMEs in the pooling area;

[0040] The deep learning model training formula is:

[0041]

[0042] Among them, L is the loss function value, yi is the true value, y^i is the predicted value, and N is the number of samples;

[0043] The parameter update formula of the gradient descent algorithm is:

[0044]

[0045] Among them, θ is the model parameter, η is the learning rate, is the gradient of the loss function with respect to the parameters;

[0046] The optimization formula of genetic algorithm optimization includes:

[0047] Fitness function: F(θ) = w1·E(θ)-w2·C(θ);

[0048] Among them, F(θ) is the fitness function, θ is the optimization parameter, E(θ) is the evaluation function of SME data analysis effect, C(θ) is the resource consumption evaluation function, w_1 and w_2 are weight coefficients;

[0049] The selection operation formula is:

[0050] Among them, P_i is the probability of an individual being selected, F_i is the fitness value of the individual, and N is the population size;

[0051] Crossover operation: θ child = r·θ parent1 +(1-r)·θ parent2

[0052] Among them, θ child is the offspring parameter, θ parent1 and θ parent2 is the parent parameter, r is a random number;

[0053] Mutation operation: θ new =θ old +Δθ;

[0054] Among them, θ new is the new parameter after mutation, θ old is the original parameter, and Δθ is the random change of the parameter.

[0055] In a preferred embodiment, the model evaluation module uses a cross-validation method to divide the data into a training set and a test set to ensure the fairness of the evaluation; the model is run on the test set to generate prediction results, which are compared with the true labels to calculate various evaluation indicators; the evaluation results are presented to data scientists through a visual report, including a confusion matrix, a ROC curve graph, and a performance indicator table; based on the evaluation results, the data scientist may decide to return to the model training module for parameter tuning or try different algorithms; once the model performance reaches expectations, it will be marked as deployable and ready to be integrated into the software application.

[0056] In a preferred embodiment, the software framework integration submodule seamlessly integrates the trained machine learning model with the existing software framework; automatically generates configuration files, service classes and controllers, and encapsulates the machine learning model as a RESTful API service.

[0057] In a preferred embodiment, the API generation and management submodule is responsible for automatically generating and managing the API interface of the machine learning model, so that the model can be called by other applications or systems in the form of a service; specific operations include defining API endpoints, parameters, request types, and return formats.

[0058] In a preferred embodiment, the deployment and monitoring submodule is responsible for deploying the integrated software application to the production environment and performing real-time monitoring to ensure the stable operation of the system; during the deployment process, the submodule is also responsible for handling the environment configuration and dependency installation details; after the deployment is completed, the monitoring function will be started to collect system performance indicators, log information and abnormal alarms in real time; when the system detects that the CPU usage rate is continuously higher than the threshold, it will automatically trigger an alarm to notify the administrator; through the deployment and monitoring submodule, the enterprise ensures the high availability and rapid response of the software application and provides stable technical support for the business.

[0059] In a preferred embodiment, the user interface and interaction module logs into the system through the user interface during use, accesses specific data analysis tools and reports according to permissions; selects the data set to be analyzed on the interface, and sets the analysis parameters; after the system receives the request, the background processes the data and generates corresponding analysis reports and visualization charts.

[0060] In a preferred embodiment, the system management and maintenance module first sets up a maintenance plan in the module, including database backup, system update and security check; the maintenance plan is set to be executed once a week and automatically started during off-peak hours; in the backup task, the module will automatically copy the data in the database to a remote server or cloud storage service; the system update task will check the versions of all software components and apply the latest security patches and functional updates; during the execution of the maintenance task, the system management and maintenance module will generate a detailed report to record the operation results and problems found; if potential security risks or performance issues are found, the system will automatically send an alert to the administrator so that timely action can be taken.

[0061] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:

[0062] 1. In the present invention, by cleaning noise and outliers and applying normalization processing, the data quality of the input model is ensured, laying a solid foundation for subsequent analysis. This not only improves the consistency of the data, but also makes the model training more stable and efficient. The use of CNN to automatically extract features avoids the subjective bias of manual selection, improves the objectivity and efficiency of feature extraction, and the output high-quality feature vectors effectively enhance the learning ability and predictive performance of the model. In the deep learning model training link, by designing a suitable model structure and applying the back propagation algorithm, the intelligent analysis of small and medium-sized enterprise data and the output of recommended parameters are realized, and the fitting ability and generalization ability of the model are significantly improved.

[0063] 2. In the present invention, the genetic algorithm is combined for optimization, the output of the deep learning model is used as the initial population, and by defining a fitness function that comprehensively considers the effect and resource consumption, multi-objective optimization is achieved, and a parameter combination that better meets actual needs is found. The model update and adaptive learning mechanism ensure that the model can learn online based on real-time feedback data, adapt to environmental changes in a timely manner, and continuously improve prediction accuracy. Therefore, the integration of the model training module not only improves the efficiency and quality of data processing, but also enhances the learning ability and prediction performance of the model, realizes multi-objective optimization and adaptive learning, and provides more intelligent and precise decision-making support and services for small and medium-sized enterprises, effectively promoting the innovation and development of small and medium-sized enterprises in the process of digital transformation.

[0064] 3. In the present invention, the software framework integration submodule seamlessly integrates various data processing, model training and artificial intelligence algorithm components into a unified software framework, thereby achieving collaborative work and efficient data flow between different modules, enabling small and medium-sized enterprises to quickly build and iterate their own applications. The API generation and management submodule automatically generates standardized and easy-to-use API interfaces, providing small and medium-sized enterprises with flexible data access and function call methods, promoting interconnection and intercommunication between the system and external systems, and providing small and medium-sized enterprises with convenient secondary development capabilities. The deployment and monitoring submodule realizes one-click deployment and real-time monitoring of the software. Through the automated deployment process, it reduces the complexity and error rate of manual operations, improves the deployment efficiency of the software, and provides reliable technical support for small and medium-sized enterprises. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 is the overall system block diagram of the present invention;

[0066] Figure 2 This is a system block diagram of the software integration and deployment module in the present invention. DETAILED DESCRIPTION

[0067] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0068] Example:

[0069] Reference Figure 1-2 , a development system for multi-source heterogeneous data-driven artificial intelligence integrated software for small and medium-sized enterprises, characterized in that: the system includes: data crawling submodule, multi-source data access submodule, data preprocessing module, model training module, model evaluation module, software integration and deployment module, user interface and interaction module and system management and maintenance module;

[0070] The software integration and deployment module is internally provided with a software framework integration submodule, an API generation and management submodule, and a deployment and monitoring submodule;

[0071] The data crawling submodule and the multi-source data access submodule serve as data entry points, collecting data from different sources such as the Internet, databases, APIs, file systems, and IoT devices, and transmitting these data to the data preprocessing module.

[0072] The data preprocessing module cleans, converts and performs feature engineering on the data to ensure data quality and consistency, preparing for subsequent model training. The processed data enters the model training module.

[0073] The model training module uses integrated machine learning and deep learning algorithms to train the model, while optimizing the model performance through the hyperparameter tuning submodule. The trained model will be moved to the model evaluation module.

[0074] The model evaluation module performs performance evaluation, visual analysis, and model interpretation to ensure the accuracy and interpretability of the model.

[0075] The software integration and deployment module integrates the model with the software framework, generates an API interface, and supports one-click deployment to the cloud platform or local server, while providing real-time monitoring capabilities.

[0076] The user interface and interaction module provides a visual interface, user management and task scheduling functions, allowing users to interact with the system conveniently.

[0077] The system management and maintenance module is responsible for system configuration, log management, security and backup, ensuring the stable operation of the system and data security.

[0078] The data crawling submodule first defines the crawling target, including a list of related websites and product information that needs to be crawled (such as price, specifications, user reviews, etc.). Next, the developer configures the crawler task in the data crawling submodule, sets the crawling frequency (such as once a day), data storage path (such as local database or cloud storage), and user agent and IP pool to avoid being blocked by the target website. After the configuration is complete, the crawler automatically starts running, accesses the target website according to preset rules, parses the web page content, extracts product information, and stores it in the specified location. Data analysts can view the crawling results in real time to ensure the accuracy and completeness of the data. Once data anomalies or crawling failures are found, the system will automatically send an alarm to notify relevant personnel so that the crawler strategy can be adjusted or the problem can be fixed in time.

[0079] Multi-source data access submodule Configure data source links in the multi-source data access submodule, including connections to relational databases, real-time data streams accessed through API interfaces, and regularly uploaded CSV files. During the configuration process, define the frequency of data extraction (such as once an hour), data conversion rules (such as unified date format, handling missing values), and target storage for data loading (such as data warehouse). After configuration, the module automatically performs the data extraction, transformation, and loading (ETL) process to integrate data from different sources into a unified data warehouse. Business analysts can use this integrated data for analysis and generate various reports and insights. If there are delays or errors in the data access process, the system management and maintenance module will record the relevant logs and notify the data engineer to troubleshoot and repair.

[0080] The data preprocessing module uses the minimum-maximum standardization method to perform data standardization. The specific steps are as follows:

[0081] S1. Calculate the maximum and minimum values: For each feature, find the maximum and minimum values ​​in the Internet security information dataset;

[0082] S2. Apply transformation formula: For each feature value x in the Internet security information dataset, use the following formula for transformation:

[0083]

[0084] Among them, x norm is the transformed value;

[0085] S3. Transform the Internet security information dataset: Apply the above transformation to each feature of the entire dataset; after minimum-maximum normalization, all features will be scaled to the range of 0 to 1.

[0086] The model training module uses the CNN algorithm to perform adaptive learning based on historical data and real-time feedback from SMEs. The specific steps include:

[0087] Data preprocessing: Clean the collected SME data to remove noise and outliers. Normalize the SME data and scale all features to the same scale.

[0088] Feature extraction: Use deep learning networks (such as convolutional neural networks (CNN)) to automatically extract the features of SME data. Output feature vectors for subsequent optimization processing.

[0089] Deep learning model training: Design a deep learning model with the extracted feature vector as input and the recommended parameters of SME data as output. Use historical SME data to train the model and optimize the model parameters through the back propagation algorithm.

[0090] Genetic algorithm optimization: The output parameters of the deep learning model are used as the initial population of the genetic algorithm. The fitness function is defined to reflect the effect and resource consumption of SME data. Through the iterative search of the genetic algorithm,

[0091] Find the optimal parameter combination.

[0092] Model update and adaptive learning: Based on real-time feedback data, the deep learning model is updated using online learning strategies, enabling the model to adapt to environmental changes and improve prediction accuracy.

[0093] The data normalization formula is:

[0094] Among them, x is the original data, x′ is the normalized data, μ is the mean of the data,

[0095] σ is the standard deviation of the data;

[0096] · The convolutional layer formula of CNN is:

[0097] · The convolutional layer formula is:

[0098]

[0099] Among them, I is the input SME data, K is the convolution kernel, and (i, j) is the position of the output SME data;

[0100] The pooling layer formula for maximum pooling is:

[0101]

[0102] Among them, P is the data feature of SMEs after pooling, and R is the data feature set of SMEs in the pooling area;

[0103] The deep learning model training formula is:

[0104]

[0105] Among them, L is the loss function value, yi is the true value, y^i is the predicted value, and N is the number of samples.

[0106] The parameter update formula of the gradient descent algorithm is:

[0107]

[0108] Among them, θ is the model parameter, η is the learning rate, is the gradient of the loss function with respect to the parameters.

[0109] The optimization formula of genetic algorithm optimization includes:

[0110] Fitness function: F(θ) = w1·E(θ)-w2·C(θ);

[0111] Among them, F(θ) is the fitness function, θ is the optimization parameter, E(θ) is the evaluation function of the data analysis effect of SMEs, C(θ) is the resource consumption evaluation function, and w_1 and w_2 are weight coefficients.

[0112] The selection operation formula is:

[0113] Among them, P_i is the probability of an individual being selected, F_i is the fitness value of the individual, and N is the population size.

[0114] Crossover operation: θ child = r·θ parent1 +(1-r)·θ parent2

[0115] Among them, θ child is the offspring parameter, θ parent1 and θ parent2 is the parent parameter and r is a random number.

[0116] Mutation operation: θ new =θ old +Δθ;

[0117] Among them, θ new is the new parameter after mutation, θ old is the original parameter, and Δθ is the random change of the parameter.

[0118] The model evaluation module uses a cross-validation method to divide the data into training and test sets to ensure the fairness of the evaluation. The model is run on the test set to generate predictions, which are compared with the true labels and various evaluation indicators are calculated. The evaluation results are presented to the data scientist through a visual report, including a confusion matrix, ROC curve graph, and performance indicator table. Based on the evaluation results, the data scientist may decide to return to the model training module for parameter tuning or try different algorithms. Once the model performance reaches the expected level, it will be marked as deployable and ready to be integrated into the software application.

[0119] The software framework integration submodule seamlessly integrates the trained machine learning model with the existing software framework. It supports a variety of popular frameworks, such as Spring Boot, Django, Flask, etc., to facilitate the rapid construction of applications. Taking Spring Boot as an example, this submodule can automatically generate configuration files, service classes, and controllers to encapsulate machine learning models as RESTful API services. In addition, it also supports custom framework integration to meet enterprise-specific needs. In this way, the model and software are closely integrated, development efficiency is improved, and the stability and scalability of the software are guaranteed.

[0120] The API generation and management submodule is responsible for automatically generating and managing the API interface of the machine learning model, so that the model can be called by other applications or systems in the form of a service. Specific operations include defining API endpoints, parameters, request types (such as POST, GET, etc.), and return formats (such as JSON, XML, etc.). At the same time, this submodule also provides functions such as API version control, permission authentication, and traffic control to ensure the security and stability of the API. Through the API generation and management submodule, enterprises can easily embed machine learning models into existing business processes to achieve intelligent upgrades.

[0121] The deployment and monitoring submodule is responsible for deploying the integrated software application to the production environment and performing real-time monitoring to ensure the stable operation of the system. Taking deployment to the cloud platform as an example, this submodule supports one-click deployment to mainstream cloud service providers such as AWS, Azure, and Alibaba Cloud, and automatically configures virtual machines, load balancing, databases and other resources. During the deployment process, the submodule is also responsible for handling details such as environment configuration and dependency installation. After the deployment is completed, the monitoring function will be started to collect system performance indicators (such as CPU usage, memory usage, network traffic, etc.), log information and abnormal alarms in real time. When the system detects that the CPU usage continues to be higher than the threshold, it will automatically trigger an alarm to notify the administrator. Through the deployment and monitoring submodule, enterprises can ensure the high availability and rapid response of software applications and provide stable technical support for the business.

[0122] The user interface and interaction module logs in to the system through the user interface during use, and accesses specific data analysis tools and reports according to permissions. Select the data set to be analyzed on the interface (such as customer purchase history, market trend analysis), and set the analysis parameters (such as time range, product category). After the system receives the request, the background processes the data and generates corresponding analysis reports and visual charts. Marketers can explore this data interactively and gain a deep understanding of customer behavior through operations such as filtering and drilling. Based on the analysis results, formulate promotion strategies, and create and schedule marketing activities in the user interface. Throughout the process, the user interface and interaction module provides real-time feedback, such as data processing progress bar, analysis result update notification, etc. In addition, the module also records user operation logs for auditing and optimization by the system management and maintenance module.

[0123] The system management and maintenance module first sets up a maintenance plan in the module, including database backup, system updates, and security checks. The maintenance plan can be set to run once a week and automatically start during off-peak hours. In the backup task, the module automatically copies the data in the database to a secure storage location, such as a remote server or cloud storage service. The system update task checks the versions of all software components and applies the latest security patches and feature updates. Security checks include scanning system vulnerabilities, verifying user permissions, and reviewing access logs. During the execution of maintenance tasks, the system management and maintenance module generates detailed reports to record the results of the operation and the problems found. If potential security risks or performance issues are found, the system automatically sends an alert to the administrator so that timely action can be taken. Through such a workflow, the system management and maintenance module ensures the continued reliability and security of the software development system.

[0124] In the present invention, by cleaning noise and outliers and applying normalization processing, the data quality of the input model is ensured, laying a solid foundation for subsequent analysis. This not only improves the consistency of the data, but also makes the model training more stable and efficient. The use of CNN to automatically extract features avoids the subjective bias of manual selection, improves the objectivity and efficiency of feature extraction, and the output high-quality feature vectors effectively enhance the learning ability and predictive performance of the model. In the deep learning model training link, by designing a suitable model structure and applying the back propagation algorithm, the intelligent analysis of small and medium-sized enterprise data and the output of recommended parameters are realized, and the fitting ability and generalization ability of the model are significantly improved.

[0125] In the present invention, the genetic algorithm is combined for optimization, the output of the deep learning model is used as the initial population, and by defining a fitness function that comprehensively considers the effect and resource consumption, multi-objective optimization is achieved, and a parameter combination that better meets actual needs is found. The model update and adaptive learning mechanism ensure that the model can learn online based on real-time feedback data, adapt to environmental changes in a timely manner, and continuously improve prediction accuracy. Therefore, the integration of the model training module not only improves the efficiency and quality of data processing, but also enhances the learning ability and prediction performance of the model, realizes multi-objective optimization and adaptive learning, and provides more intelligent and precise decision support and services for small and medium-sized enterprises, effectively promoting the innovation and development of small and medium-sized enterprises in the process of digital transformation.

[0126] In the present invention, the software framework integration submodule realizes the collaborative work and efficient data flow between different modules by seamlessly integrating various data processing, model training and artificial intelligence algorithm components into a unified software framework, so that small and medium-sized enterprises can quickly build and iterate their own applications. The API generation and management submodule automatically generates standardized and easy-to-use API interfaces, provides flexible data access and function call methods for small and medium-sized enterprises, promotes the interconnection and intercommunication between the system and external systems, and provides small and medium-sized enterprises with convenient secondary development capabilities. The deployment and monitoring submodule realizes one-click deployment and real-time monitoring of the software. Through the automated deployment process, it reduces the complexity and error rate of manual operations, improves the deployment efficiency of the software, and provides reliable technical support for small and medium-sized enterprises.

[0127] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprises" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0128] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A development system for multi-source heterogeneous data-driven artificial intelligence integrated software for small and medium-sized enterprises, characterized by: The system includes: a data crawling submodule, a multi-source data access submodule, a data preprocessing module, a model training module, a model evaluation module, a software integration and deployment module, a user interface and interaction module, and a system management and maintenance module; The software integration and deployment module is internally provided with a software framework integration submodule, an API generation and management submodule, and a deployment and monitoring submodule; The data crawling submodule and the multi-source data access submodule serve as data entry points to collect data from different sources such as the Internet, databases, APIs, file systems, and IoT devices, and transmit the data to the data preprocessing module; The data preprocessing module cleans, converts and performs feature engineering on the data to ensure the quality and consistency of the data and prepare for subsequent model training. The processed data enters the model training module. The model training module uses integrated machine learning and deep learning algorithms to train the model, and optimizes the model performance through the hyperparameter tuning submodule; the trained model will be moved to the model evaluation module. The model evaluation module performs performance evaluation, visualization analysis, and model interpretation to ensure the accuracy and interpretability of the model; The software integration and deployment module integrates the model with the software framework, generates an API interface, supports one-click deployment to the cloud platform or local server, and provides real-time monitoring capabilities; The user interface and interaction module provides a visual interface, user management and task scheduling functions, allowing users to interact with the system conveniently; The system management and maintenance module is responsible for system configuration, log management, security and backup to ensure stable operation of the system and data security.

2. The multi-source heterogeneous data-driven artificial intelligence integrated software development system for small and medium-sized enterprises as claimed in claim 1, characterized in that: The data crawling submodule first defines the crawling target, including a list of related websites and product information to be crawled; then, the developer configures the crawler task in the data crawling submodule, sets the crawling frequency, data storage path, and user agent and IP pool to avoid being blocked by the target website; After the configuration is completed, the crawler automatically starts running, accesses the target website according to the preset rules, parses the webpage content, extracts product information, and stores it in the specified location; data analysts view the crawling results in real time to ensure the accuracy and completeness of the data; Once data anomalies or crawling failures are detected, the system will automatically send an alert to notify relevant personnel so that crawler strategies can be adjusted or problems can be fixed in a timely manner. The multi-source data access submodule configures data source links in the multi-source data access submodule, including connections to relational databases, real-time data streams accessed through API interfaces, and regularly uploaded CSV files; during the configuration process, the frequency of data extraction, data conversion rules, and target storage for data loading are defined; Once configured, the module automatically performs the data extraction, conversion, and loading process, integrating data from different sources into a unified data warehouse; Business analysts use this integrated data for analysis and generate various reports and insights; if there are delays or errors in the data access process, the system management and maintenance module will record the relevant logs and notify the data engineer to troubleshoot and repair.

3. The development system of multi-source heterogeneous data-driven artificial intelligence integrated software for small and medium-sized enterprises as claimed in claim 1, characterized in that: The data preprocessing module uses the minimum-maximum standardization method to perform data standardization processing, and the specific steps are as follows: S1. Calculate the maximum and minimum values: For each feature, find the maximum and minimum values ​​in the Internet security information dataset; S2. Apply transformation formula: For each feature value x in the Internet security information dataset, use the following formula for transformation: Among them, x norm is the transformed value; S3. Transform the Internet security information dataset: Apply the above transformation to each feature of the entire dataset; after minimum-maximum normalization, all features will be scaled to the range of 0 to 1.

4. The multi-source heterogeneous data-driven artificial intelligence integrated software development system for small and medium-sized enterprises as claimed in claim 1, characterized in that: The model training module uses the CNN algorithm to perform adaptive learning based on historical data and real-time feedback from SMEs, specifically including the following steps: Data preprocessing: clean the collected SME data to remove noise and outliers; normalize the SME data and scale all features to the same scale; Feature extraction: Use deep learning networks to automatically extract features of SME data; output feature vectors for subsequent optimization processing; Deep learning model training: Design a deep learning model with the extracted feature vector as input and the recommended parameters of SME data as output; use historical SME data to train the model and optimize the model parameters through the back propagation algorithm; Genetic algorithm optimization: Use the output parameters of the deep learning model as the initial population of the genetic algorithm; define the fitness function to reflect the effect and resource consumption of SME data; find the optimal parameter combination through iterative search of the genetic algorithm; Model update and adaptive learning: Use online learning strategies to update deep learning models based on real-time feedback data, enabling the model to adapt to environmental changes and improve prediction accuracy; The data normalization formula is: Among them, x is the original data, x′ is the normalized data, μ is the mean of the data, and σ is the standard deviation of the data; · The convolutional layer formula of CNN is: · The convolutional layer formula is: Among them, I is the input SME data, K is the convolution kernel, and (i, j) is the position of the output SME data; The pooling layer formula for maximum pooling is: Among them, P is the data feature of SMEs after pooling, and R is the data feature set of SMEs in the pooling area; The deep learning model training formula is: Among them, L is the loss function value, yi is the true value, y^i is the predicted value, and N is the number of samples; The parameter update formula of the gradient descent algorithm is: Among them, θ is the model parameter, η is the learning rate, is the gradient of the loss function with respect to the parameters; The optimization formula of genetic algorithm optimization includes: Fitness function: F(θ) = w1·E(θ)-w2·C(θ); Among them, F(θ) is the fitness function, θ is the optimization parameter, E(θ) is the evaluation function of SME data analysis effect, C(θ) is the resource consumption evaluation function, w_1 and w_2 are weight coefficients; The selection operation formula is: Among them, P_i is the probability of an individual being selected, F_i is the fitness value of the individual, and N is the population size; Crossover operation: θ child = r·θ parent1 +(1-r)·θ parent2 Among them, θ child is the offspring parameter, θ parent1 and θ parent2 is the parent parameter, r is a random number; Mutation operation: θ new =θ old +Δθ; Among them, θ new is the new parameter after mutation, θ old is the original parameter, and Δθ is the random change of the parameter.

5. The multi-source heterogeneous data-driven artificial intelligence integrated software development system for small and medium-sized enterprises as claimed in claim 1, characterized in that: The model evaluation module uses a cross-validation method to divide the data into a training set and a test set to ensure the fairness of the evaluation; The model is run on the test set to generate predictions, which are compared with the true labels and various evaluation indicators are calculated. The evaluation results are presented to data scientists through visual reports, including confusion matrices, ROC curves, and performance indicator tables. Based on the evaluation results, data scientists may decide to return to the model training module for parameter tuning or try different algorithms. Once the model performance meets expectations, it will be marked as deployable and ready to be integrated into the software application.

6. The multi-source heterogeneous data-driven artificial intelligence integrated software development system for small and medium-sized enterprises as claimed in claim 1, characterized in that: The software framework integration submodule seamlessly integrates the trained machine learning model with the existing software framework; automatically generates configuration files, service classes and controllers, and encapsulates the machine learning model as a RESTful API service.

7. The multi-source heterogeneous data-driven artificial intelligence integrated software development system for small and medium-sized enterprises as claimed in claim 1, characterized in that: The API generation and management submodule is responsible for automatically generating and managing the API interface of the machine learning model, so that the model can be called by other applications or systems in the form of a service; specific operations include defining API endpoints, parameters, request types, and return formats.

8. The multi-source heterogeneous data-driven artificial intelligence integrated software development system for small and medium-sized enterprises as claimed in claim 1, characterized in that: The deployment and monitoring submodule is responsible for deploying the integrated software application to the production environment and performing real-time monitoring to ensure the stable operation of the system; during the deployment process, the submodule is also responsible for handling the environment configuration and dependency installation details; After deployment is complete, the monitoring function will be started to collect system performance indicators, log information and abnormal alarms in real time; when the system detects that the CPU usage rate is continuously higher than the threshold, it will automatically trigger an alarm to notify the administrator; through the deployment and monitoring sub-modules, enterprises ensure the high availability and rapid response of software applications, and provide stable technical support for the business.

9. The multi-source heterogeneous data-driven artificial intelligence integrated software development system for small and medium-sized enterprises as claimed in claim 1, characterized in that: During use, the user interface and interaction module logs into the system through the user interface, accesses specific data analysis tools and reports according to permissions; selects the data set to be analyzed on the interface, and sets the analysis parameters; after the system receives the request, the background processes the data and generates corresponding analysis reports and visualization charts.

10. The multi-source heterogeneous data-driven artificial intelligence integrated software development system for small and medium-sized enterprises as claimed in claim 1, characterized in that: The system management and maintenance module first sets up a maintenance plan in the module, including database backup, system update and security check; the maintenance plan is set to be executed once a week and automatically started during off-peak hours; in the backup task, the module will automatically copy the data in the database to a remote server or cloud storage service; the system update task will check the versions of all software components and apply the latest security patches and functional updates; during the execution of the maintenance task, the system management and maintenance module will generate a detailed report to record the operation results and problems found; if potential security risks or performance issues are found, the system will automatically send an alert to the administrator so that timely action can be taken.

Citation Information

Patent Citations

  • Real-time decision support system based on machine learning

    CN118886728A

  • Data analysis system based on artificial intelligence

    CN119066423A

  • Cross-cloud resource scheduling method and system based on deep learning

    CN119299519A