Intelligent decision-making method and system for job progress based on big data
Through real-time collection, cleaning and segmented analysis of production site data, combined with historical records optimization scheduling strategies, the problem of real-time fusion of production site data is solved, and intelligent management and continuous optimization of the production process are achieved.
Patent Information
- Application Number
- CN202510398496.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-04
AI Technical Summary
It is difficult for the existing technology to realize real-time collection and fusion analysis of multi-source heterogeneous data such as equipment operation status, personnel dynamics and process progress at the production site, resulting in frequent deviations from the plan and affecting production efficiency.
By obtaining real-time data flow at the production site, uniform format processing and cleaning are performed, sliding window segment analysis is used to generate early warning signals and scheduling instructions, and dynamic comparison and feedback mechanisms are combined with historical records to optimize the scheduling strategy to build a dual closed-loop control process.
It realizes timely detection of abnormal situations such as equipment failure, personnel flow and material shortage, generates corresponding early warning signals and scheduling instructions, optimizes the production process, and improves production efficiency and resource utilization.
Smart Images

Figure CN120258234A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular, to an intelligent decision-making method and system for job progress based on big data. Background Art
[0002] Intelligent decision-making for job progress based on big data is a core research direction in the field of modern industrial production management. Its importance lies in significantly improving production efficiency, optimizing resource allocation, and coping with complex and changing production environments through data-driven means. Research in this field is directly related to whether an enterprise can maintain agility and cost advantages in the fierce market competition, and is also a key technical support for promoting the transformation and upgrading of intelligent manufacturing. With the wide application of industrial Internet and big data technologies, the intelligent management of job progress has become an irreversible development trend.
[0003] However, existing methods have shown obvious limitations when dealing with actual production scenarios. Traditional scheduling decision-making systems often rely on static models and manual experience, and it is difficult to adapt to the dynamic changes in the production site. Even if some data analysis technologies are introduced, their processing capabilities are still limited by the incompleteness of data collection, the lag of analysis, and the lack of real-time decision-making, resulting in frequent deviations of job progress from the plan and affecting the overall production efficiency.
[0004] Focusing on the core challenges, real-time control has become the primary technical bottleneck. The production site involves the real-time collection and fusion analysis of multi-source heterogeneous data such as equipment operating status, personnel dynamics, and process progress. However, interference factors in complex environments, such as equipment failures and material shortages, make it difficult to achieve millisecond-level perception and scheduling adjustment. In addition, data processing faces a dilemma: edge computing is limited by resources and cannot support complex decision-making; although cloud computing has sufficient computing power, it is difficult to meet the real-time requirements due to network latency. At the same time, the construction of a global optimization scheduling model for multiple workshops and multiple processes is complex, and the rapid solution of the optimal solution in a dynamic environment has become another key problem. Therefore, how to coordinate the computing task allocation between the edge and the cloud in a complex and changing production environment, achieve the real-time collection, fusion, and analysis of multi-source data, and build an efficient global optimization scheduling model to ensure the seamless connection between job progress and the plan has become a key issue in the field of intelligent decision-making for job progress based on big data. Summary of the Invention
[0005] The present invention provides an intelligent scheduling method for job progress based on big data, mainly including: Obtain the real-time data stream of the production site, where the real-time data stream includes equipment operation parameters, personnel location coordinates, and process progress indicators. Perform unified format processing on the real-time data stream through a standardized protocol conversion module, and store the processed data in the local database; extract the uniformly formatted data from the local database, perform segmented processing on the data using a fixed-duration sliding window, and perform extreme value removal, missing value filling, and outlier filtering operations on each data segment in sequence according to the preset cleaning rules to generate cleaned data segments; perform aggregated batch processing on the cleaned data segments. If it is detected that the equipment operation parameters exceed the preset fault threshold, generate an equipment fault warning signal and a corresponding scheduling instruction template. If it is detected that the change range of the personnel location coordinates exceeds the movement threshold, trigger a personnel reassignment instruction. If the process progress indicator is lower than the preset value of the material, generate a replenishment warning signal; package the warning signal and the scheduling instruction template, combine the historical production record set, dynamically calculate the real-time state weight using a preset algorithm, and generate an input factor set including time dimension and space dimension through multi-dimensional data fusion; input the input factor set into a pre-designed calculation framework, perform in-depth analysis processing based on the global model variables, and generate a structured scheduling plan according to the optimization objective function; perform dynamic comparison between the structured scheduling plan and the current real-time data stream. If it is detected that the deviation value exceeds the preset threshold, extract the deviation data set and feedback it to the cloud platform to trigger the model update mechanism; adjust the real-time state weight parameters according to the deviation data set, update the global model variables, and regenerate the scheduling plan to form a double-closed-loop control process including a local execution layer and a cloud optimization layer; when the scheduling plan passes the verification, execute the operation instructions in the scheduling instruction template, and at the same time collect the changed data after execution, and upload the changed data to the cloud platform through the incremental update mechanism.
[0006] The technical solution provided by the embodiment of the present invention may include the following beneficial effects: By collecting the equipment operation parameters, personnel location, and process progress data of the production site in real time, performing standardized processing and cleaning on the data, and performing segmented analysis based on the sliding window. The present invention can detect abnormal situations such as equipment failures, personnel flow, and material shortages in a timely manner, and generate corresponding warning signals and scheduling instructions. By integrating historical production records and real-time status, the present invention also continuously optimizes the scheduling strategy through dynamic comparison and feedback mechanisms, so as to realize the intelligent management and continuous optimization of the production process, and improve production efficiency and resource utilization rate. Description of the Drawings
[0007] Figure 1 It is a flowchart of an intelligent decision-making method for operation progress based on big data according to the present invention.
[0008] Figure 2 It is a specific flowchart of step S104 of the present invention.
[0009] Figure 3 This is the specific flowchart of step S105 of the present invention. Detailed implementation manner
[0010] To further understand the content of the present invention, the present invention will be described in detail with reference to the accompanying drawings and embodiments. The present application will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the sake of description, only the parts related to the invention are shown in the drawings.
[0011] Such as Figures 1-3 , a method for intelligent decision-making on operation progress based on big data in this embodiment, the method specifically includes: Step S101, obtain the real-time data stream of the production site, the real-time data stream includes equipment operation parameters, personnel position coordinates and process progress indicators, perform format unification processing on the real-time data stream through a standardized protocol conversion module, and store the processed data in a local database.
[0012] Obtain the original data set, perform format unification processing on the original data set to obtain a format-unified data set; for the format-unified data set, perform integrity verification, if missing values are detected, fill them through a preset method to obtain a complete data set; perform anomaly detection based on the complete data set, perform correlation analysis on the abnormal data points and the personnel position coordinates to obtain an abnormal distribution data set; according to the abnormal distribution data set, store the data with unified format and no anomalies in the database to obtain a stored complete data record.
[0013] Specifically, for the acquisition and format unification processing of the original data set, assume that the operation scenario is the personnel positioning data management in a factory workshop. The original data set may include information such as the positioning coordinates, timestamps, and employee numbers of employees, but these data may come from different devices and have different formats.
[0014] For example, the time format recorded by some devices is "2023-10-01 08:00:00", while that of others is "20231001T080000". In the format unification processing, all time formats will be converted to the standard form, such as "2023-10-01 08:00:00", the coordinate data will be unified into the two-dimensional plane coordinate form, such as (x, y), and ensure that the employee number is a string of unified length. This can avoid errors caused by inconsistent formats in subsequent processing and improve the efficiency and accuracy of data processing. Perform integrity verification on the format-unified data set. Assume that it is found during verification that some data is missing timestamps or coordinate values.
[0015] For example, a certain record only has an employee number and an x - coordinate, while the y - coordinate and timestamp are empty. In one possible implementation, for missing values, they can be filled through a preset method. The interpolation method based on time series can be adopted to infer the missing timestamp by combining the previous and subsequent records. For missing coordinates, the possible y - coordinate values can be estimated according to the continuity of the employee's movement trajectory. For example, the average value of the coordinates at the previous and next moments can be taken. This method can effectively fill the data gaps, ensure the integrity of the data set, and provide a reliable basis for subsequent analysis. In the anomaly detection and correlation analysis stage, assuming that after analyzing the complete data set, it is found that an employee's coordinate changes abnormally in a short period of time, such as jumping instantaneously from (10, 20) to (100, 200), far exceeding the normal movement speed.
[0016] In one embodiment, anomalies are detected by setting a speed threshold (such as 10 meters per second), and these anomaly points are associated with the employee's position coordinates to generate an anomaly distribution data set.
[0017] For example, it is found through analysis that the anomaly points are concentrated near a certain device, which may indicate that the positioning error is caused by device signal interference. This correlation analysis helps to locate the root cause of the problem and improve the system stability.
[0018] Preferably, for the anomaly distribution data set, the data without anomalies is stored in the database.
[0019] For example, the uniformly formatted and anomaly - free employee location data is stored in the database according to the time series to form a stored complete data record. This storage method is convenient for subsequent query and analysis, such as counting the daily activity trajectories of employees or optimizing the workshop layout.
[0020] It can be understood that after removing the abnormal data, the records in the database are more accurate, avoiding analysis biases caused by abnormal values and providing reliable support for management decisions.
[0021] It should be noted that through the above process, the data is normalized, completed, and made accurate from collection to storage, which not only improves the data quality but also lays a foundation for subsequent job analysis.
[0022] For example, based on the stored complete data record, the work efficiency of employees can be further analyzed or potential safety hazards can be discovered, significantly improving the operation efficiency of the factory.
[0023] Step S102: Extract the uniformly formatted data from the local database, segment the data using a fixed - duration sliding window, and sequentially perform extreme value removal, missing value filling, and outlier filtering operations on each data segment according to preset cleaning rules to generate the cleaned data segments.
[0024] Obtain a data set with a unified format from the local database, and use a fixed-duration sliding window to segment the data set to obtain a set of data segments; for each data segment in the set of data segments, remove extreme values according to the preset 3σ rule, use linear interpolation to fill in missing values, and filter out outliers according to the preset box plot rule to obtain a set of cleaned data segments; extract the stationary feature fields of each data segment from the set of cleaned data segments, and use the K-means clustering algorithm to group the stationary feature fields to obtain a set of grouped feature data.
[0025] Specifically, after obtaining a data set with a unified format from the local database, using a fixed-duration sliding window for segmentation processing can cut a continuous data stream into data segments for multiple time periods.
[0026] For example, in the scenario of personnel location data management in a factory workshop, assume that the database stores the coordinate and timestamp information of employees, and the sliding window is set to one segment every 5 minutes, and the data of one day is divided into 288 segments. This segmentation method is convenient for subsequent analysis of the behavior characteristics of employees in different time periods.
[0027] It can be understood that a fixed-duration window can effectively capture the change trend in a short time and provide a basis for subsequent cleaning and grouping. Applying the 3σ rule to remove extreme values for each segment in the set of data segments is a common statistical method.
[0028] Exemplarily, assume that the average value of the employee coordinate data in a certain segment is (50, 60), and the standard deviation is 10. If a coordinate value is (85, 90), which exceeds the 3-fold standard deviation range, it is regarded as an extreme value and is excluded.
[0029] In a possible implementation manner, if it is found that a data point is missing after exclusion, for example, the coordinate record of an employee is missing within 5 minutes, linear interpolation can be used to fill it.
[0030] Specifically, based on the coordinates of the previous moment (40, 50) and the next moment (60, 70), the value of the missing point is inferred to be (50, 60). This method can maintain the continuity of the data.
[0031] It should be noted that the box plot rule is used to further filter out outliers and can identify outlier points beyond the normal range.
[0032] For example, the y coordinate values of employees in a certain segment are mostly between 50 and 70, but a record is 200, which far exceeds the upper limit of the box plot and is filtered out. This double cleaning mechanism ensures the reliability of the data segments.
[0033] Preferably, the set of data segments after cleaning is more suitable for extracting stable feature fields. For example, calculate the average value and variance of the employee coordinates within each segment as an index to describe the stability of the employee's position. After extracting the stable feature fields from the set of data segments after cleaning, using the K-means clustering algorithm for grouping is an unsupervised learning method.
[0034] In one embodiment, assuming that the extracted features include the average coordinates and moving speed, and setting the value of K to 3, employees can be divided into a "stationary group", a "slow moving group", and a "fast moving group".
[0035] For example, if the average coordinates of an employee change by less than 5 meters within 5 minutes and the speed is 0.1 m / s, the employee is classified into the "stationary group"; if the speed of another employee is 2 m / s, the employee is classified into the "slow moving group".
[0036] Specifically, the clustering results can reflect the activity patterns of employees in the workshop.
[0037] It can be understood that this grouping helps to identify the distribution of employees in different working states.
[0038] In one possible implementation, the set of grouped feature data after clustering can be used to analyze the gathering areas of employees in the workshop.
[0039] For example, if the average coordinates of a group of employees are concentrated near (30, 40), it may indicate that this is the main workstation. This kind of analysis can provide a basis for optimizing personnel scheduling.
[0040] Preferably, the set of grouped feature data can also reveal the time patterns of employees' activities. For example, the "fast moving group" appears more frequently during shift change periods.
[0041] It should be noted that this method transforms the original data into more meaningful structured information through feature extraction and clustering, laying a foundation for subsequent management decisions.
[0042] Step S103: Perform aggregation processing on the set of data segments after cleaning to generate an aggregated data set, and generate a warning signal and a scheduling instruction according to the aggregated data set.
[0043] In one possible implementation, the data segments after cleaning usually contain various information in the factory workshop. By generating an aggregated data set through batch aggregation processing, the scattered data can be integrated into more valuable structured data for analysis.
[0044] For example, in a certain manufacturing factory, the aggregated data set may include equipment operation parameters such as temperature and pressure, personnel position coordinates such as the real-time position of employees in the workshop, and process progress indicators such as the percentage of tasks completed on a certain production line.
[0045] Specifically, the batch aggregation process can summarize the cleaned data segments within each hour on an hourly basis to generate an aggregated data set containing the above three types of metrics.
[0046] Taking the processing of device operation parameters as an example, if the parameter value exceeds the preset fault threshold, a device fault warning signal is generated. Exemplarily, assume that the normal operating temperature range of a certain device is 20 to 50 degrees Celsius, and the preset fault threshold is 55 degrees Celsius. If the device temperature reaches 60 degrees Celsius within a certain hour as shown in the aggregated data set, the system will generate a fault warning signal. The process of generating a scheduling instruction by associating the warning signal with a preset scheduling instruction template may be to extract a corresponding plan template for high-temperature faults from the database, such as reducing the device load or pausing operation.
[0047] In one embodiment, if the warning signal indicates that the device temperature is too high, the scheduling instruction template may include an adjustment parameter field, such as increasing the device fan speed from 1000 revolutions per minute to 1500 revolutions per minute.
[0048] Step S104, pack the warning signal and the scheduling instruction into data, and generate a preliminary scheduling plan in combination with the historical record; Obtain the packed data set. For the warning signal fields contained therein, use a logistic regression algorithm to classify the warning signals to obtain the classification result of the warning signals; if the classification result belongs to a high priority level, then according to the classification result, in combination with the historical production record set, extract the corresponding adjustment fields from the scheduling instruction template to generate a preliminary adjustment instruction set; for the preliminary adjustment instruction set, use the K-means clustering algorithm to group the instruction fields to obtain the grouped instruction subsets; according to the grouped instruction subsets, in combination with the real-time state weight, if the real-time state weight exceeds the preset threshold, then perform a priority sorting on the instruction subsets to generate a sorted instruction sequence; extract the time dimension information from the sorted instruction sequence, and adjust the instruction execution order according to the time dimension information to obtain an optimized scheduling plan; obtain the optimized scheduling plan, in combination with the space dimension information, and divide the scheduling execution range according to the space dimension information to generate a preliminary scheduling plan.
[0049] In a possible implementation manner, after obtaining the packed data set, first use a logistic regression algorithm to classify the warning signal fields contained therein. The core of the logistic regression algorithm lies in performing weighted analysis on data features through a probability model to judge the priority of the warning signal.
[0050] For example, in a certain manufacturing factory, the warning signal fields may include categories such as equipment overload, insufficient materials, and abnormal personnel flow. Suppose the characteristic value of a certain equipment overload signal is relatively high. After logistic regression analysis, its probability value is 0.85, exceeding the high-priority threshold of 0.8, and it is classified as a high-priority signal.
[0051] Specifically, if the classification result is high-priority, then in combination with the historical production record set, extract the adjustment fields from the scheduling instruction template to generate a preliminary adjustment instruction set. The historical production record set may contain adjustment plans in similar overload situations in the past year, such as reducing equipment power or increasing maintenance frequency.
[0052] Exemplarily, if the historical record shows that a certain equipment returned to normal after reducing its power by 10%, the scheduling instruction template will extract the "power adjustment" field to generate a preliminary adjustment instruction set, such as "reduce power by 10%" and "increase cooling time by 20 minutes".
[0053] Preferably, use the K-means clustering algorithm to group the preliminary adjustment instruction set to obtain instruction subsets. The goal of K-means clustering is to cluster similar instruction fields into one category.
[0054] For example, "reduce power by 10%" and "reduce power by 15%" may be grouped into one group, while "increase cooling time" is grouped into another group. After grouping, if the real-time state weight, such as the current equipment load rate of 0.9, exceeds the preset threshold of 0.7, then perform priority sorting on the instruction subsets. Suppose the "reduce power" group has a higher priority, and the sorted instruction sequence is "reduce power by 10% - increase cooling time by 20 minutes".
[0055] It can be understood that adjust the instruction execution order according to the time dimension information to generate an optimized instruction execution plan. The time dimension information may include the peak operation period of the equipment. Suppose the peak period is from 8:00 to 10:00 in the morning, then the optimization plan will give priority to arranging the "reduce power" instruction to be executed before the peak. Further combining the space dimension information, such as the workshop area where the equipment is located, divide the instruction execution scope into production lines A and B. The final plan may be "production line A executes power adjustment first, and production line B executes cooling later".
[0056] In one embodiment, generate an execution log including time and space dimensions according to the final plan, such as "8:00 - production line A - reduce power by 10%", and store it in the historical production record set. The stored log can provide data support for subsequent analysis.
[0057] Step S105, the preliminary scheduling plan is dynamically compared with the dynamic input parameter sequence. If it is detected that the deviation value exceeds the preset threshold, then optimize the scheduling plan based on the deviation value to obtain the final optimized scheduling plan.
[0058] Obtain the real-time data of the sensor interface to generate a dynamic input parameter sequence; determine whether the matching degree between the dynamic input parameter sequence and the scheduling scheme exceeds a preset threshold. If it exceeds the threshold, extract the deviation data set that exceeds the threshold; for the deviation data set, use the K-means clustering algorithm for feature analysis to obtain an updated parameter set; adjust the matching fields in the scheduling scheme according to the updated parameter set to generate an optimized scheme version; determine whether the matching degree between the optimized scheme version and the data stream reaches a preset standard. If it reaches the preset standard, determine the optimized scheme version as the final scheme.
[0059] Specifically, obtaining the real-time data of the sensor interface to generate a dynamic input parameter sequence, the core lies in capturing real-time changes in the production environment. For example, in a certain manufacturing factory, sensors may monitor equipment temperature, vibration frequency, and power consumption.
[0060] Exemplarily, the temperature data is updated once per second, forming a sequence such as 25°C, 26°C, 27°C, and the vibration frequency may be 120Hz, 122Hz, 121Hz. This dynamic sequence reflects the immediate state of equipment operation and provides a basis for subsequent analysis.
[0061] It can be understood that the real-time nature of the data ensures that the parameter sequence can respond promptly to production fluctuations. Determining whether the matching degree between the dynamic input parameter sequence and the scheduling scheme exceeds a preset threshold is crucial for evaluating whether the existing scheme is suitable for the current state.
[0062] In a possible implementation, assume that the scheduling scheme requires the temperature not to exceed 26°C, while the real-time sequence shows 27°C, exceeding the threshold by 1°C. Similarly, if the vibration frequency threshold is 120Hz, 122Hz in the sequence also exceeds the setting.
[0063] Specifically, the deviation data set may be extracted as a temperature deviation of 1°C and a vibration frequency deviation of 2Hz. This judgment method intuitively reflects the gap between the scheme and the actual operation. For the deviation data set, use the K-means clustering algorithm for feature analysis, aiming to discover the laws behind the deviations.
[0064] It should be noted that K-means groups the deviation data to identify potential patterns. For example, a temperature deviation of 1°C and a vibration deviation of 2Hz may be grouped into one category, indicating that equipment overheating is related to increased vibration. In another embodiment, if power consumption deviation is also included in the analysis, the clustering may be divided into two groups: one dominated by temperature and vibration, and the other dominated by power fluctuations. The updated parameter set may be adjusted to a temperature upper limit of 27°C and a vibration frequency of 123Hz, reflecting the new features after clustering. Adjusting the matching fields in the scheduling scheme according to the updated parameter set to generate an optimized scheme version emphasizes the targeted adjustment of the fields.
[0065] Preferably, if the upper temperature limit is adjusted to 27°C, the matching field may change from "reduce the cooling fan speed" to "maintain the current fan speed".
[0066] Exemplarily, after the vibration frequency is adjusted, the field may become "reduce the load by 5%".
[0067] In one embodiment, the adjusted solution may be "08:00 - Device B - maintain the fan speed, load reduced by 5%". This adjustment ensures that the solution better fits the real-time data. Judging whether the matching degree of the optimized solution version and the data stream reaches the preset standard determines the final confirmation of the solution.
[0068] Specifically, if the preset standard is a matching degree of 95%, after the new solution runs, the temperature stabilizes at 26.5°C, the vibration drops to 121Hz, and the matching degree may reach 96%. For example, through simulation runs, it is found that the vibration further stabilizes after the load is reduced, and the matching degree continuously remains above the standard.
[0069] It can be understood that this verification method confirms the feasibility of the solution.
[0070] In one embodiment, if the matching degree does not meet the standard, the field can be further fine-tuned, such as "reduce the load by 7%", until the requirements are met. The final scheduling solution such as "08:10 - Device B - load reduced by 7%" is then determined as the optimal scheduling version.
[0071] In the embodiments of the present application, by collecting the device operation parameters, personnel positions, and process progress data at the production site in real time, the data is standardized and cleaned, and segmented analysis is performed based on a sliding window. The present invention can detect abnormal situations such as equipment failures, personnel flow, and material shortages in a timely manner, and generate corresponding warning signals and scheduling instructions. By integrating historical production records and real-time status, the present invention also continuously optimizes the scheduling strategy through a dynamic comparison and feedback mechanism, thereby realizing the intelligent management and continuous optimization of the production process, and improving production efficiency and resource utilization rate.
[0072] The embodiments of the present application also provide an intelligent decision-making system for job progress based on big data. The system includes: The real-time data stream acquisition module at the operation site acquires the real-time data stream at the operation site. The real-time data stream includes equipment operation parameters, personnel location coordinates, and process progress indicators. The standardized protocol conversion module performs format unification processing on the real-time data stream and stores the processed data in the local database; the data cleaning module extracts the data with unified format from the local database, performs segmented processing on the data using a fixed-duration sliding window, and sequentially performs extreme value removal, missing value filling, and outlier filtering operations on each data segment according to the preset cleaning rules to generate the cleaned data segments; the data aggregation module performs aggregation processing on the set of the cleaned data segments to generate an aggregated data set, and generates an early warning signal and a scheduling instruction according to the aggregated data set; the preliminary scheduling plan generation module packs the early warning signal and the scheduling instruction, and generates a preliminary scheduling plan in combination with the historical record; the optimized scheduling plan generation module dynamically compares the preliminary scheduling plan with the dynamic input parameter sequence. If it is detected that the deviation value exceeds the preset threshold, the scheduling plan is optimized based on the deviation value to obtain the final optimized scheduling plan.
[0073] The embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the data security control method based on the hybrid cloud as described above.
[0074] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0075] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0076] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes Figure 1 and / or boxes Figure 1 specified in one or more of the processes and / or boxes.
[0077] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the processes Figure 1 and / or boxes Figure 1 specified in one or more of the boxes.
[0078] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0079] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.
[0080] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0081] It should also be noted that the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0082] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. An intelligent decision-making method for job progress based on big data, characterized in that, The method includes: Obtaining the real-time data stream at the job site, where the real-time data stream includes equipment operation parameters, personnel location coordinates, and process progress indicators, performing unified format processing on the real-time data stream through a standardized protocol conversion module, and storing the processed data in a local database; extracting the uniformly formatted data from the local database, performing segmented processing on the data using a fixed-duration sliding window, and sequentially performing extreme value removal, missing value filling, and outlier filtering operations on each data segment according to preset cleaning rules to generate cleaned data segments; performing aggregation processing on the set of cleaned data segments to generate an aggregated data set, generating warning signals and scheduling instructions based on the aggregated data set; packing the warning signals and scheduling instructions into data, and generating a preliminary scheduling plan in combination with historical records; dynamically comparing the preliminary scheduling plan with a dynamic input parameter sequence, and if a deviation value exceeding a preset threshold is detected, optimizing the scheduling plan based on the deviation value to obtain a final optimized scheduling plan.
2. The method according to claim 1, characterized in that, The obtaining of the real-time data stream at the job site, where the real-time data stream includes equipment operation parameters, personnel location coordinates, and process progress indicators, performing unified format processing on the real-time data stream through a standardized protocol conversion module, and storing the processed data in a local database includes: Obtaining an original data set, performing unified format processing on the original data set to obtain a uniformly formatted data set; Performing integrity verification on the uniformly formatted data set, and if missing values are detected, filling them through a preset method to obtain a complete data set; Performing anomaly detection based on the complete data set, associating anomaly data points with personnel location coordinates for correlation analysis to obtain an anomaly distribution data set; Based on the anomaly distribution data set, storing the data with unified format and no anomalies in the database.
3. The method according to claim 1, wherein The extracting of the uniformly formatted data from the local database, performing segmented processing on the data using a fixed-duration sliding window, and sequentially performing extreme value removal, missing value filling, and outlier filtering operations on each data segment according to preset cleaning rules to generate cleaned data segments includes: Obtaining the uniformly formatted data set in the local database, performing segmented processing on the data set using a fixed-duration sliding window to obtain a set of data segments; For each data segment in the set of data segments, removing extreme values according to the preset 3σ rule, filling missing values using the linear interpolation method, and filtering outliers according to the preset box plot rule to obtain a set of cleaned data segments.
4. The method according to claim 1, wherein The packing of the warning signals and scheduling instructions into data, and generating a preliminary scheduling plan in combination with historical records includes: Obtaining a packed data set, classifying the warning signals using a logistic regression algorithm for the warning signal fields included therein to obtain a classification result of the warning signals; If the classification result belongs to a high priority level, then according to the classification result, in combination with the historical production record set, extracting the corresponding adjustment fields from the scheduling instruction template to generate a set of preliminary adjustment instructions; For the set of preliminary adjustment instructions, the K-means clustering algorithm is used to group the instruction fields to obtain the grouped instruction subsets; According to the grouped instruction subsets and combined with the real-time status weight, if the real-time status weight exceeds the preset threshold, the instruction subsets are sorted by priority to generate a sorted instruction sequence; Extract the time dimension information from the sorted instruction sequence, and adjust the instruction execution order according to the time dimension information to obtain an optimized scheduling plan; Obtain the optimized scheduling plan, combine the space dimension information, and divide the scheduling execution scope according to the space dimension information to generate a preliminary scheduling plan.
5. The method according to claim 1, characterized in that, The preliminary scheduling plan is dynamically compared with the dynamic input parameter sequence. If it is detected that the deviation value exceeds the preset threshold, the scheduling plan is optimized based on the deviation value to obtain the final optimized scheduling plan, including: Obtain the real-time data of the sensor interface and generate a dynamic input parameter sequence; Judge whether the matching degree between the dynamic input parameter sequence and the scheduling plan exceeds the preset threshold. If it exceeds the threshold, extract the deviation data set that exceeds the threshold; For the deviation data set, perform feature analysis using the K-means clustering algorithm to obtain an updated parameter set; Adjust the matching fields in the preliminary scheduling plan according to the updated parameter set to generate an optimized scheduling plan; Judge whether the matching degree between the optimized scheduling plan and the dynamic input parameter sequence reaches the preset standard. If it reaches the preset standard, determine the optimized scheduling plan as the optimized generated scheduling plan.
6. An intelligent decision-making system for job progress based on big data, characterized in that, The system includes: A real-time data stream acquisition module at the job site, which acquires the real-time data stream at the job site. The real-time data stream includes equipment operation parameters, personnel position coordinates, and process progress indicators. The real-time data stream is uniformly processed in format through a standardized protocol conversion module, and the processed data is stored in a local database; a data cleaning module, which extracts uniformly formatted data from the local database, performs segmented processing on the data using a fixed-duration sliding window, and sequentially performs extreme value removal, missing value filling, and outlier filtering operations on each data segment according to preset cleaning rules to generate cleaned data segments; a data aggregation module, which performs aggregation processing on the set of cleaned data segments to generate an aggregated data set, and generates a warning signal and a scheduling instruction according to the aggregated data set; a preliminary scheduling plan generation module, which packs the warning signal and the scheduling instruction, and generates a preliminary scheduling plan in combination with historical records; an optimized scheduling plan generation module, which dynamically compares the preliminary scheduling plan with the dynamic input parameter sequence. If it is detected that the deviation value exceeds the preset threshold, the scheduling plan is optimized based on the deviation value to obtain the final optimized scheduling plan.
7. The system according to claim 6, wherein The real-time data stream acquisition module at the job site includes: Obtain the original data set, and perform unified format processing on the original data set to obtain a uniformly formatted data set; For the uniformly formatted data set, perform integrity verification. If missing values are detected, fill them through a preset method to obtain a complete data set; Anomaly detection is performed based on the complete data set, and the anomaly data points are associated and analyzed with the personnel location coordinates to obtain an anomaly distribution data set; Based on the anomaly distribution data set, the data with unified format and no anomalies is stored in the database.
8. The system according to claim 6, wherein The data cleaning module includes: Obtain a data set with a unified format in the local database, and segment the data set using a fixed-duration sliding window to obtain a set of data segments; For each data segment in the set of data segments, remove extreme values according to the preset 3σ rule, fill in missing values using linear interpolation, and filter out anomaly values according to the preset box plot rule to obtain a set of cleaned data segments.
9. The system according to claim 6, characterized in that, The preliminary scheduling plan generation module includes: Obtain the packaged data set, and classify the warning signals using a logistic regression algorithm for the warning signal fields included therein to obtain the classification results of the warning signals; If the classification result belongs to the high priority level, then according to the classification result, combined with the historical production record set, extract the corresponding adjustment fields from the scheduling instruction template to generate a set of preliminary adjustment instructions; For the set of preliminary adjustment instructions, use the K-means clustering algorithm to group the instruction fields to obtain a grouped instruction subset; According to the grouped instruction subset, combined with the real-time status weight, if the real-time status weight exceeds the preset threshold, then perform priority sorting on the instruction subset to generate a sorted instruction sequence; Extract the time dimension information from the sorted instruction sequence, and adjust the instruction execution order according to the time dimension information to obtain an optimized scheduling plan; Obtain the optimized scheduling plan, combined with the space dimension information, and divide the scheduling execution scope according to the space dimension information to generate a preliminary scheduling plan.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the intelligent decision-making method for job progress based on big data provided in any one of claims 1-5.
Citation Information
Cited By
Intelligent control method and system for workpiece heat treatment equipment
CN121232764A
Workshop operation state real-time monitoring and control method and system
CN121279609A
Automatic production data protection method and system for semiconductor manufacturing equipment
CN121413008A
Methods and systems for automated production data protection in semiconductor manufacturing equipment
CN121413008B