Port cost tracing method based on multi-source data fusion

By constructing a composite identifier for work units and modeling the distribution of Dirichlet evidence, combined with an improved Aumann–Shapley method, the problem of temporal and spatial alignment of multi-source data in port cost tracing was solved, achieving refined cost allocation and tracing, and improving the accuracy and interpretability of the results.

CN121809843APending Publication Date: 2026-04-07YANTAI PORT GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing port cost tracing methods lack the ability to integrate multi-source data, making it difficult to achieve time alignment and spatial mapping. This results in coarse cost accounting granularity, which cannot meet the needs of refined tracing. Furthermore, the lack of precise binding of equipment energy consumption, staff hours, and financial vouchers leads to distorted allocation results.

Method used

By constructing a composite identifier for work units, introducing Dirichlet evidence distribution modeling, and combining it with the improved Aumann–Shapley method, we can achieve temporal and spatial alignment of multi-source data. Furthermore, by adjusting the marginal contribution through confidence and conflict levels, we can perform fine-grained allocation and generate reliable fusion results.

Benefits of technology

It achieves high-precision cost traceability at the port operation unit level, outputs cost results with multi-dimensional distribution and factor contribution, ensures the accuracy and interpretability of traceability results, and reduces the deviation rate in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121809843A_ABST
    Figure CN121809843A_ABST
Patent Text Reader

Abstract

The invention discloses a port cost tracing method based on multi-source data fusion, and the method comprises the steps: S1, collecting port operation multi-source data, and carrying out the standardization processing; s2, executing time alignment and space alignment; s3, generating an operation unit composite identifier, constructing an operation unit model, and forming an operation unit data set; s4, performing evidence modeling and fusion through Dirichlet evidence distribution, and generating a fusion result with confidence and conflict degree; s5, generating a power curve for the energy consumption data in the fusion result, and performing integration to obtain energy consumption; processing the personnel man-hour record to obtain a man-hour amount; performing text extraction on the financial voucher record to generate cost information; s6, identifying a cost driving factor, and generating a job unit apportionment value set by adopting an improved Aumann-Shapley method; and S7, calculating the total cost of the operation unit, and outputting a cost tracing result. According to the method, the precision and interpretability of port operation cost tracing are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multi-source data fusion and cost traceability calculation technology, and in particular to a port cost traceability method based on multi-source data fusion. Background Technology

[0002] With the continuous improvement of port automation and intelligent management requirements, multi-dimensional cost accounting, including energy consumption, labor costs, and depreciation, in port operations has gradually become an important foundation for enterprise management and decision-making. Existing cost accounting methods are mostly based on the summary of account period expenses in the financial system, or rely on single-dimensional data from equipment management systems and operational statistics systems for estimation. However, in complex port operation scenarios, the following problems are common: existing methods often lack the ability to integrate multi-source heterogeneous data; data formats such as ship plans, equipment operation records, personnel work hours, and financial vouchers differ significantly, with inconsistent time bases and spatial identifiers, making it difficult to form a unified data system; there is a lack of precise time alignment and spatial mapping between multi-source data; equipment energy consumption curves, personnel work hour records, and expense vouchers cannot be accurately bound at the operational unit level, resulting in coarse accounting granularity, only able to allocate costs at the team or account period level, lacking the ability to finely trace individual containers or instructions.

[0003] Most existing technologies rely on traditional average allocation or simple weighting methods, which fail to reflect the varying contributions of different equipment, personnel, and operating conditions to costs. For example, the energy consumption curves of quay cranes and yard cranes show significant differences in peak values ​​during operating periods, but these differences are not effectively reflected in traditional allocation models, leading to distorted allocation results. Overlapping labor hours across days and shift handovers are not effectively handled, often causing errors in work hour statistics. Unstructured information from financial voucher images is not effectively integrated with operational unit data, resulting in a disconnect between cost collection and actual operational processes. Furthermore, existing cost allocation methods are mostly static models, lacking consideration of dynamic factors such as operational routes, time-of-day electricity prices, and spatial location differences, thus failing to meet the accurate accounting needs of a dynamic port operating environment.

[0004] In terms of data fusion and reliability processing, existing methods typically employ simple data averaging or weighted fusion, failing to consider reliability differences and conflicting relationships between evidence from different sources. For records involving various uncertainties and contradictions, such as energy consumption, working hours, and costs, traditional methods often lack quantitative modeling of evidence confidence and conflict levels, leading to insufficient reliability of the fusion results. Furthermore, the lack of correction mechanisms for temporal overlap and spatial matching rates during multi-source data alignment further exacerbates the impact of data inconsistencies on the results. These shortcomings result in existing port cost traceability systems often only achieving coarse-grained comparisons of total accounts when faced with large-scale, multi-dimensional data, making it difficult to support the needs of refined cost management and anomaly traceability at the operational unit level.

[0005] Therefore, how to provide a port cost tracing method based on multi-source data fusion is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] This invention addresses the issues of temporal inconsistency, missing spatial identifiers, and conflicts in multi-source data for port operations by constructing a composite identifier for work units and introducing Dirichlet evidence distribution modeling. This achieves high-precision alignment of ship plans, equipment energy consumption, personnel man-hours, and financial documents. The improved Aumann–Shapley method incorporates confidence and conflict levels to adjust marginal contributions, and combines time-of-use electricity pricing and demand charges to achieve fine allocation of energy consumption, labor, and cost factors. Ultimately, this enables work unit cost traceability and factor contribution decomposition, improving the accuracy and interpretability of the results.

[0007] A port cost tracing method based on multi-source data fusion according to an embodiment of the present invention includes the following steps: S1. Collect multi-source data related to port operations, perform standardization processing, and output a standardized dataset; S2. Perform time and spatial alignment on the standardized dataset, and map the location information to the yard grid and berth index to obtain aligned data; S3. Generate a composite identifier for the work unit based on the alignment data and construct a work unit model. Map the equipment energy consumption records, personnel work hour records and financial vouchers to the corresponding work units to form a work unit dataset. S4. Introduce the Dirichlet evidence distribution into the job unit dataset, perform evidence modeling and fusion, and generate a fusion result with confidence and conflict. S5. Generate a power curve from the energy consumption data in the fusion result, and integrate the power curve during the working period to obtain the energy consumption; process the personnel work hour records to obtain the working hours; perform text extraction on the financial voucher records to generate cost information; S6. Based on the energy consumption, working hours and cost information, identify cost driving factors and use the improved Aumann-Shapley method to generate a set of work unit allocation values; S7. Calculate the total cost of the work unit based on the energy consumption, working hours and allocated value set, and output the cost traceability result.

[0008] Optionally, step S1 includes: S11. Collect records of ship plans, berthing times, berth numbers, container numbers, work instructions, process identifiers, yard location records, and equipment allocation instructions in the operation management system; S12. Collect operational status data, start and stop times, number of lifting operations, handling distance, equipment energy consumption records, and operation logs of quay cranes, yard cranes, reach stackers, forklifts, automated guided vehicles, and container trucks. S13. Collect personnel working hours records, including port entry and exit times, shift schedules, job assignments, single container working hours, and team working hours; S14. Collect financial voucher records, including expense vouchers, invoice records, energy costs, labor costs, expense allocation records, and total costs for the payment period; S15. Perform uniform format conversion processing on the collected data, standardize the timestamp, location information, and identifier fields, and output a standardized dataset.

[0009] Optionally, step S2 includes: S21. Establish a unified time benchmark for the ship plans, berthing times, work instructions, equipment operation records, personnel working hours records and expense voucher records in the standardized dataset; S22. Based on the unified time benchmark, the granularity of the timestamps of multi-source data is adjusted, the minute-level, hour-level and day-level records are converted into a consistent time interval, and placeholder marks are filled for missing time periods, and overlapping time periods are merged in chronological order. S23. Map the berth number and berthing time according to the unified time base to generate a berth index and form a ship operation space code; S24. Map the storage yard location records according to the unified time base to generate a storage yard grid code and form a storage space code; S25. Combine the equipment number, personnel number and expense item number based on the unified time base and spatial code to generate a data index that can be used to uniquely locate the work unit; S26. Output the aligned dataset, which includes a unified time series, berth index, yard grid code, equipment number, personnel number, cost item number, and work unit index.

[0010] Optionally, step S3 includes: S31. Extract the container number, process identifier, equipment number, personnel number, operation period, berth index and yard grid code from the aligned dataset to generate a composite identifier that can uniquely locate the operation unit. S32. Map the device energy consumption records according to the composite identifier to generate an energy consumption subset; S33. Map the personnel work hour records according to the composite identifier to generate a subset of work hours; S34. Map the financial voucher records according to the composite identifier to generate a cost subset; S35. Based on the composite identifier, the energy consumption subset, working hour subset, and cost subset are jointly bound in the time dimension, space dimension, resource dimension, and cost dimension to form a work unit dataset containing time, space, equipment, personnel, and cost information.

[0011] Optionally, step S4 includes: S41. Extract equipment energy consumption records, personnel working hours records and financial voucher records from the work unit dataset, divide the equipment energy consumption status interval, working hour status interval and cost status interval, and establish an evidence set. S42. Generate a Dirichlet evidence distribution model for each type of evidence, record the parameter vectors of each state, and label the reliability weights of the evidence sources. S43. Perform prior subtraction and normalization on the parameter vectors of each evidence source, calculate the support quantity of each state, and obtain the belief value sequence and uncertainty value of each type of evidence. S44. Based on the reliability weight, the parameter vectors and belief value sequences of each evidence source are weighted and combined to generate the fused Dirichlet distribution and fused belief sequence. S45. According to the preset benchmark rate, the fused belief sequence and uncertainty are matched and synthesized, and the confidence level of each state is calculated. S46. Statistically analyze the dispersion of the belief value sequence of each evidence source for each state to obtain the inter-source dispersion sequence; calculate the deviation measure between each evidence source and the fused belief sequence to obtain the deviation sequence; synthesize the conflict degree according to the preset weight, and introduce the time overlap rate and spatial matching rate penalty coefficient for correction to generate the conflict degree of each state. S47. Output a set of evidence fusion results containing the confidence and conflict levels of each state.

[0012] Optionally, step S5 includes: S51. Extract equipment energy consumption records, personnel work hour records, and financial voucher records from the evidence fusion results based on the composite identifier of the work unit. S52. Resample the energy consumption records of the equipment at preset time intervals, perform outlier removal and smoothing processing, generate a power sequence for the working period and form a power curve. S53. The power curve is truncated and summed according to the working period to generate energy consumption, and the peak, normal and valley periods are marked. S54. Perform cross-day segmentation, overlapping time period clipping, and shift handover marking on personnel work hour records, and aggregate work hours by work period to generate work hour quantity; S55. Perform text extraction on financial voucher records, extract expense items, amounts, voucher numbers and accounting periods, match them according to the composite identifier of the work unit, and generate expense information; S56. The energy consumption, working hours and cost information are bound to the composite identifier of the work unit, and attached to the confidence and conflict degree generated during the evidence fusion process to form a work unit feature set.

[0013] Optionally, step S6 includes: S61. Based on the composite identifier of the work unit, gather energy consumption, working hours and cost information, and combine it with the confidence and conflict degree generated during the evidence fusion process to form a set of candidate factors; S62. Identify energy consumption factors, labor factors and cost factors in the candidate factor set respectively, and establish a cost-driving factor sequence. S63. Generate a sequential operation path based on the operation instructions, and divide the time into peak hours, normal hours and valley hours, while performing spatial segmentation in combination with the berth index and yard grid code; S64. In the improved Aumann–Shapley method, confidence is introduced as a weighting coefficient and conflict degree is introduced as a penalty factor to adjust the marginal contribution of each cost driver in different time and space segments, and generate a phased allocation value. S65. Introduce time-of-use pricing and demand charges into the calculation of phased apportionment value, and implement peak responsibility allocation and capacity apportionment for public energy consumption and costs; S66. Aggregate the phased allocation values ​​by work unit to form the allocation results of energy consumption, labor and expenses, and maintain a corresponding relationship with the total cost over the billing period; S67. Perform consistency verification and boundary truncation on the allocation results, and output the allocation value set of the job unit.

[0014] Optionally, step S7 includes: S71. Summarize the energy consumption, working hours and allocated value based on the composite identifier of the work unit, and calculate the total cost of the work unit; S72. Combine the total cost with the berth index and the yard grid code to perform spatial mapping and generate a multi-dimensional cost distribution map of the port area. S73. Decompose the total cost according to the energy consumption factor, labor factor and expense factor, and output the factor contribution sequence; S74. Compare the factor contribution sequence with the total cost of the payment period to generate a difference test value; S75. When the difference test value exceeds the threshold, re-execute the time base correction and spatial coding calibration. S76. Re-output the updated total cost and factor contribution sequence, and generate a cost traceability result set containing cost distribution, factor contribution, and difference test values; S77. The cost traceability result set is stored in a structured manner.

[0015] The beneficial effects of this invention are: This invention addresses the issues of inconsistent time bases, missing spatial identifiers, and data conflicts in port operations' multi-source heterogeneous data at the acquisition end by constructing a composite identifier for work units and introducing Dirichlet evidence distribution modeling and fusion. It employs unified time base correction and yard grid and berth index mapping, combined with time granularity adjustment and spatial coding generation, to achieve high-precision alignment and standardization of ship plans, equipment energy consumption records, personnel man-hour records, and financial voucher records at the work unit level. In the evidence fusion stage, a multi-state Dirichlet distribution model is established, jointly introducing the reliability weight of evidence sources with time overlap rate and spatial matching rate penalty factors to dynamically calculate the confidence and conflict degree of each state, outputting a fusion result with quantifiable reliability. In the cost allocation stage, an improved Aumann–Shapley method is proposed, introducing confidence degree as a weighting coefficient and conflict degree as a penalty factor into the factor marginal contribution calculation. Combined with time-of-use electricity pricing and demand-based pricing mechanisms, it achieves refined allocation of energy consumption, labor, and cost factors across different time periods and spatial units, avoiding the distortion caused by the average allocation of traditional methods. Ultimately, it achieves total cost traceability at the port operation unit level, outputting cost results that include multi-dimensional distribution and factor contributions. Through dynamic comparison and correction with total cost over the payment period, it ensures the accuracy and interpretability of the traceability results. Attached Figure Description

[0016] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0017] Figure 1 This is a schematic diagram of the overall process of a port cost tracing method based on multi-source data fusion proposed in this invention; Figure 2 This is a schematic diagram of the evidence modeling and fusion process based on Dirichlet evidence distribution in this invention; Figure 3 This is a schematic diagram illustrating the cost driver allocation of the improved Aumann–Shapley method in this invention; Detailed Implementation

[0018] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0019] refer to Figure 1-3 A port cost tracing method based on multi-source data fusion includes the following steps: S1. Collect multi-source data related to port operations, perform standardization processing, and output a standardized dataset; S2. Perform time and spatial alignment on the standardized dataset, and map the location information to the yard grid and berth index to obtain aligned data; S3. Generate a composite identifier for the work unit based on the alignment data and construct a work unit model. Map the equipment energy consumption records, personnel work hour records and financial vouchers to the corresponding work units to form a work unit dataset. S4. Introduce the Dirichlet evidence distribution into the job unit dataset, perform evidence modeling and fusion, and generate a fusion result with confidence and conflict. S5. Generate a power curve from the energy consumption data in the fusion result, and integrate the power curve during the working period to obtain the energy consumption; process the personnel work hour records to obtain the working hours; perform text extraction on the financial voucher records to generate cost information; S6. Based on the energy consumption, working hours and cost information, identify cost driving factors and use the improved Aumann-Shapley method to generate a set of work unit allocation values; S7. Calculate the total cost of the work unit based on the energy consumption, working hours and allocated value set, output the cost traceability result, and perform factor contribution decomposition on the traceability result, compare it with the total over the billing period, and perform data realignment and parameter recalibration when the difference exceeds the threshold.

[0020] In this embodiment, step S1 includes: S11. Collect records of ship plans, berthing times, berth numbers, container numbers, work instructions, process identifiers, yard location records, and equipment allocation instructions in the operation management system; S12. Collect operational status data, start and stop times, number of lifting operations, handling distance, equipment energy consumption records, and operation logs of quay cranes, yard cranes, reach stackers, forklifts, automated guided vehicles, and container trucks. S13. Collect personnel working hours records, including port entry and exit times, shift schedules, job assignments, single container working hours, and team working hours; S14. Collect financial voucher records, including expense vouchers, invoice records, energy costs, labor costs, expense allocation records, and total costs for the payment period; S15. Perform uniform format conversion processing on the collected data, standardize the timestamp, location information, and identifier fields, and output a standardized dataset.

[0021] This invention comprehensively collects multi-source data related to port operations, covering management data such as vessel plans, berth information, container number records, work instructions, and process identifiers. It also collects the operating status, energy consumption records, and operation logs of equipment such as quay cranes, yard cranes, reach stackers, forklifts, and container trucks, combining this with personnel hour records and financial vouchers to ensure coverage of energy consumption, labor, and cost dimensions. Simultaneously, it performs a unified format conversion on the collected data, standardizing timestamps, location fields, and identification information to form a standardized dataset that can be used for subsequent time and spatial alignment, providing a consistent data foundation for cost traceability.

[0022] In this embodiment, step S2 includes: S21. Establish a unified time benchmark for the ship plans, berthing times, work instructions, equipment operation records, personnel working hours records and expense voucher records in the standardized dataset; S22. Based on the unified time benchmark, the granularity of the timestamps of multi-source data is adjusted, the minute-level, hour-level and day-level records are converted into a consistent time interval, and placeholder marks are filled for missing time periods, and overlapping time periods are merged in chronological order. S23. Map the berth number and berthing time according to the unified time base to generate a berth index and form a ship operation space code; S24. Map the storage yard location records according to the unified time base to generate a storage yard grid code and form a storage space code; S25. Combine the equipment number, personnel number and expense item number based on the unified time base and spatial code to generate a data index that can be used to uniquely locate the work unit; S26. Output the aligned dataset, which includes a unified time series, berth index, yard grid code, equipment number, personnel number, cost item number, and work unit index.

[0023] This invention achieves time alignment of multi-source data by establishing a unified time benchmark and unifies the granularity of minute-level, hourly-level, and daily-level records. It also fills in missing segments and merges overlapping segments to ensure data continuity and consistency. At the spatial level, it maps berth numbers to berthing times to generate berth indexes and generates grid codes for yard locations, forming refined spatial positioning. Furthermore, it binds equipment numbers, personnel numbers, and expense item numbers with time and spatial codes to generate data indexes that uniquely locate work units, achieving multi-dimensional fusion indexing and providing an innovative data organization method for subsequent composite identifier generation and work unit modeling.

[0024] In this embodiment, step S3 includes: S31. Extract the container number, process identifier, equipment number, personnel number, operation period, berth index and yard grid code from the aligned dataset to generate a composite identifier that can uniquely locate the operation unit. S32. Map the device energy consumption records according to the composite identifier to generate an energy consumption subset; S33. Map the personnel work hour records according to the composite identifier to generate a subset of work hours; S34. Map the financial voucher records according to the composite identifier to generate a cost subset; S35. Based on the composite identifier, the energy consumption subset, working hour subset, and cost subset are jointly bound in the time dimension, space dimension, resource dimension, and cost dimension to form a work unit dataset containing time, space, equipment, personnel, and cost information.

[0025] In this invention, the composite identifier for a work unit is generated by combining the container number, process identifier, work period, berth index, yard grid code, equipment number, personnel number, and expense item number. The container number uniquely corresponds to the container; the process identifier distinguishes between loading / unloading, storage, or relocation stages; the work period ensures uniqueness in the time dimension; the berth index and yard grid code provide spatial positioning; the equipment number and personnel number identify the work resource; and the expense item number is associated with financial data. Each field is uniformly encoded and then concatenated in a fixed order or indexed using a database composite primary key to form a unique identifier. Based on this composite identifier, equipment energy consumption records, personnel work hour records, and financial voucher records can be accurately mapped to the corresponding work unit, achieving integrated traceability of cross-source data.

[0026] In this embodiment, step S4 includes: S41. Extract equipment energy consumption records, personnel working hours records and financial voucher records from the work unit dataset, divide the equipment energy consumption status interval, working hour status interval and cost status interval, and establish an evidence set. S42. Generate a Dirichlet evidence distribution model for each type of evidence, record the parameter vectors of each state, and label the reliability weights of the evidence sources. S43. Perform prior subtraction and normalization on the parameter vectors of each evidence source, calculate the support quantity of each state, and obtain the belief value sequence and uncertainty value of each type of evidence. S44. Based on the reliability weight, the parameter vectors and belief value sequences of each evidence source are weighted and combined to generate the fused Dirichlet distribution and fused belief sequence. S45. According to the preset benchmark rate, the fused belief sequence and uncertainty are matched and synthesized, and the confidence level of each state is calculated. S46. Statistically analyze the dispersion of the belief value sequence of each evidence source for each state to obtain the inter-source dispersion sequence; calculate the deviation measure between each evidence source and the fused belief sequence to obtain the deviation sequence; synthesize the conflict degree according to the preset weight, and introduce the time overlap rate and spatial matching rate penalty coefficient for correction to generate the conflict degree of each state. The deviation: ; in, This indicates the fusion bias, measuring the overall difference between each source of evidence and the fusion result; Indicates the total number of sources of evidence; Indicates the source index of evidence From 1 to Summing each one individually; Indicates the first The belief vector of a source of evidence in all states; The penalty coefficient: ; in, Represents the spacetime penalty term, with values... The larger the value, the worse the spatiotemporal consistency. Represents the time overlap rate, with values ​​ranging from 1 to 2. The intersection-union ratio is calculated based on the work period: intersection duration divided by union duration; Represents spatial matching rate, with values ​​ranging from 1 to 2. Calculated by dividing the berth index by the yard grid consistency ratio; The degree of conflict: ; in, Indicates the final degree of conflict; Indicates the dispersion weight; Indicates the inter-source dispersion; This represents the deviation weight, and is related to... The sum equals 1; Indicates the fusion deviation; Indicates the spacetime penalty term; S47. Output a set of evidence fusion results containing the confidence and conflict levels of each state.

[0027] In this invention, to achieve reliable fusion of multi-source heterogeneous data, evidence sets are established for equipment energy consumption records, personnel work hour records, and financial voucher records in the work unit dataset, and a state space is defined for each type of evidence: equipment energy consumption is divided by power range, personnel work hour by work hour range, and financial voucher by expense amount range. For each evidence set, a Dirichlet evidence distribution model is constructed, and the support of each state is described by a parameter vector. To improve fusion accuracy, a source reliability factor is introduced to distinguish the credibility differences between automatically collected data and manually entered data. Furthermore, during the evidence fusion process, weighted combinations are applied to various Dirichlet distribution parameters, and a conflict adjustment factor is introduced to weaken the impact of highly conflicting evidence. Through this method, a fused Dirichlet distribution is obtained, and the confidence and conflict degree of each state are calculated within this distribution. The confidence level characterizes the credibility of multi-source evidence under consistency, and the conflict degree quantifies the differences between evidence. Finally, a fusion result set containing confidence and conflict degree labels is generated, providing a basis for cost driver factor identification and allocation calculation.

[0028] In this embodiment, step S5 includes: S51. Extract equipment energy consumption records, personnel work hour records, and financial voucher records from the evidence fusion results based on the composite identifier of the work unit. S52. Resample the energy consumption records of the equipment at preset time intervals, perform outlier removal and smoothing processing, generate a power sequence for the working period and form a power curve. S53. The power curve is truncated and summed according to the working period to generate energy consumption, and the peak, normal and valley periods are marked. S54. Perform cross-day segmentation, overlapping time period clipping, and shift handover marking on personnel work hour records, and aggregate work hours by work period to generate work hour quantity; S55. Perform text extraction on financial voucher records, extract expense items, amounts, voucher numbers and accounting periods, match them according to the composite identifier of the work unit, and generate expense information; S56. The energy consumption, working hours and cost information are bound to the composite identifier of the work unit, and attached to the confidence and conflict degree generated during the evidence fusion process to form a work unit feature set.

[0029] In this invention, equipment energy consumption records, personnel work hours records, and financial voucher records are extracted from the evidence fusion results based on the work unit composite identifier. The equipment energy consumption records are resampled at fixed time intervals and outliers are removed to generate power curves for work periods and sum them up to obtain the energy consumption. The personnel work hours records are segmented across days and overlapped and clipped, and aggregated by work period to obtain the work hours. The financial voucher records are subjected to character recognition and field parsing to extract expense items and amounts to generate expense information. Finally, the energy consumption, work hours, and expense information are bound to the work unit composite identifier, and the confidence and conflict levels generated during the evidence fusion process are added to form a work unit feature set.

[0030] In this embodiment, step S6 includes: S61. Based on the composite identifier of the work unit, gather energy consumption, working hours and cost information, and combine it with the confidence and conflict degree generated during the evidence fusion process to form a set of candidate factors; S62. Identify energy consumption factors, labor factors and cost factors in the candidate factor set respectively, and establish a cost-driving factor sequence. S63. Generate a sequential operation path based on the operation instructions, and divide the time into peak hours, normal hours and valley hours, while performing spatial segmentation in combination with the berth index and yard grid code; S64. In the improved Aumann–Shapley method, confidence is introduced as a weighting coefficient and conflict degree is introduced as a penalty factor to adjust the marginal contribution of each cost driver in different time and space segments, and generate a phased allocation value. S65. Introduce time-of-use pricing and demand charges into the calculation of phased apportionment value, and implement peak responsibility allocation and capacity apportionment for public energy consumption and costs; S66. Aggregate the phased allocation values ​​by work unit to form the allocation results of energy consumption, labor and expenses, and maintain a corresponding relationship with the total cost over the billing period; S67. Perform consistency verification and boundary truncation on the allocation results, and output the allocation value set of the job unit.

[0031] In this invention, energy consumption, working hours, and cost information are aggregated based on the composite identifier of the work unit, and a candidate factor set is formed by combining the confidence and conflict degrees in the evidence fusion results. Energy consumption factors, labor factors, cost factors, and material factors are identified through classification, constructing a cost-driving factor sequence. Subsequently, a sequential work path is established, and each factor is located by combining time and spatial segmentation. In the allocation calculation, the improved Aumann–Shapley method introduces confidence as a weighting coefficient and conflict degree as a penalty factor, and combines time-of-use electricity pricing and demand charge corrections to ensure that each cost-driving factor receives a more reasonable marginal contribution allocation in the shared cost allocation, thereby forming a work unit allocation value set that conforms to the actual port operations.

[0032] In this embodiment, step S7 includes: S71. Summarize the energy consumption, working hours and allocated value based on the composite identifier of the work unit, and calculate the total cost of the work unit; S72. Combine the total cost with the berth index and the yard grid code to perform spatial mapping and generate a multi-dimensional cost distribution map of the port area. S73. Decompose the total cost according to the energy consumption factor, labor factor and expense factor, and output the factor contribution sequence; S74. Compare the factor contribution sequence with the total cost of the payment period to generate a difference test value; S75. When the difference test value exceeds the threshold, re-execute the time base correction and spatial coding calibration. S76. Re-output the updated total cost and factor contribution sequence, and generate a cost traceability result set containing cost distribution, factor contribution, and difference test values; S77. The cost traceability result set is stored in a structured manner.

[0033] This invention utilizes composite identifiers for work units to summarize energy consumption, labor hours, and allocated values, forming a total unit cost. It then combines berth indexes and yard grid codes to generate a multi-dimensional spatial distribution map. In the cost decomposition stage, the total cost is broken down into energy consumption, labor, and expense factors, outputting a sequence of factor contributions, and comparing it with the total cost over the payment period. When the difference exceeds a threshold, a recalibration of the time base and spatial code is triggered for dynamic correction. The final output is a result set containing cost distribution, factor contributions, and difference verification values, which is then stored in a structured manner.

[0034] Example 1: To verify the feasibility of this invention in practice, it was applied to the daily production management of a large container port. This port handles approximately 28,000 TEUs of containers daily, with operations encompassing ship berthing, loading and unloading, yard storage, vehicle dispatching, and financial settlement. It involves over 600 pieces of equipment and more than 1,200 personnel. In existing technologies, this port faces significant problems in cost traceability: firstly, the sampling frequencies of multi-source data differ (e.g., equipment energy consumption records are at the minute level, financial vouchers are monthly summaries, and personnel work hours are recorded at the shift level), leading to difficulties in time alignment; secondly, the spatial mapping of operations is inconsistent, lacking a direct correspondence between berths and yard locations, making it difficult to accurately allocate costs; furthermore, discrepancies exist between the total cost over the financial period and the energy consumption and work hour allocation of the operational units, often exceeding 15%, making it difficult to meet the needs of refined management.

[0035] In this embodiment, the method of the present invention was used in a typical operational cycle of the port in early August. First, the system collected vessel plans, berthing times, operational instructions, equipment operation logs, energy consumption data, personnel hours, and financial vouchers, and generated a standardized dataset through unified format conversion. During the data alignment phase, a unified time base was used to unify the granularity of minute-level equipment data, hourly-level work hour data, and daily-level financial data. Missing night shift work hours were filled with placeholder markers, and duplicate equipment start-up and shutdown events were merged in chronological order to ensure temporal continuity. At the spatial level, berth numbers and berthing times were mapped to berth indices, and yard location records were used to generate yard grid codes, forming a refined operational spatial positioning.

[0036] During the work unit construction phase, information such as container number, process, equipment number, personnel number, and work period was extracted to generate a composite identifier. For example, the identifier for a work unit consists of "berth A3-yard C5-equipment Q07-process unloading-container number SZXU1234567-time period 08:00~12:00". Under this identifier, energy consumption records are mapped to generate an energy consumption subset, personnel work hours are generated to generate a work hour subset, and financial vouchers are generated to generate an expense subset. These are then linked through multiple dimensions to form a work unit dataset.

[0037] In the evidence modeling and fusion process, this invention introduces the Dirichlet evidence distribution. Taking the unloading operation unit as an example, there is a difference between the belief value of the equipment energy consumption subset during peak power periods and the belief value of the personnel work hours subset during shift transition periods. The deviation degree is calculated through inter-source dispersion, and further penalty coefficients for time overlap rate and spatial matching rate are introduced to generate the conflict degree. The results show that in the alignment process of the energy consumption of equipment Q07 and the work hours of personnel shifts, the fusion confidence level reaches 0.91, and the conflict degree is only 0.07, effectively solving the problem of inconsistency among multi-source data.

[0038] This invention introduces an improved Aumann-Shapley method in the cost allocation process. First, confidence level is used as a weighting coefficient in the marginal contribution calculation, ensuring that factors with high data source reliability have greater weight in the allocation and reducing interference from low-confidence data. Second, conflict level is introduced as a penalty factor, automatically suppressing abnormal contributions when there are significant differences between multi-source data, thus avoiding incorrect allocation. Third, time-of-use pricing and demand charges are introduced to allocate peak-hour energy consumption additionally and off-peak-hour energy consumption less, achieving precise time-of-use cost allocation. Fourth, peak responsibility sharing and capacity allocation are adopted in the allocation of public energy consumption and depreciation, ensuring reasonable allocation of depreciation for large equipment under high load, rather than average allocation. To further verify the beneficial effects of the present invention, three sets of comparative experiments were conducted, and the experimental results are shown in Table 1: Table 1. Comparison of the method of the present invention and the traditional method under different interference scenarios.

[0039] As shown in Table 1, the method of this invention exhibits significant advantages in all three types of port operation scenarios. The deviation rate between the cost allocation result and financial expenses remains stable at around 3%, which is significantly higher than the deviation rate of over 15% for traditional methods, effectively improving the allocation accuracy. During peak periods when energy consumption and labor hours overlap, the method of this invention corrects marginal contributions by introducing confidence and conflict levels, ensuring a reasonable allocation of public energy consumption and depreciation, and keeping the difference between the allocation result and the total cost over the payment period within 3%. Especially in complex operating environments spanning multiple time periods and berths, this invention uses an improved Aumann-Shapley method combined with time-of-use electricity pricing and demand charges to dynamically adjust the contribution of energy consumption and labor, making the cost allocation result closer to actual consumption. This effectively solves the shortcomings of traditional methods in ignoring spatial segmentation and evidence conflict, significantly enhancing the accuracy and robustness of the traceability results.

[0040] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A port cost tracing method based on multi-source data fusion, characterized in that, Includes the following steps: S1. Collect multi-source data on port operations, perform standardization processing, and output a standardized dataset; S2. Perform time and spatial alignment on the standardized dataset, and map the location information to the yard grid and berth index to obtain aligned data; S3. Generate a composite identifier for the work unit based on the alignment data and construct a work unit model. Map the equipment energy consumption records, personnel work hour records and financial vouchers to the corresponding work units to form a work unit dataset. S4. Introduce the Dirichlet evidence distribution into the job unit dataset, perform evidence modeling and fusion, and generate a fusion result with confidence and conflict. S5. Generate a power curve from the energy consumption data in the fusion result, and integrate the power curve during the working period to obtain the energy consumption. Process personnel work hour records to obtain work hours; perform text extraction on financial voucher records to generate expense information; S6. Based on the energy consumption, working hours and cost information, identify cost driving factors and use the improved Aumann-Shapley method to generate a set of work unit allocation values; S7. Calculate the total cost of the work unit based on the energy consumption, working hours and allocated value set, and output the cost traceability result.

2. The port cost tracing method based on multi-source data fusion according to claim 1, characterized in that, Step S1 includes: S11. Collect records of ship plans, berthing times, berth numbers, container numbers, work instructions, process identifiers, yard location records, and equipment allocation instructions in the operation management system; S12. Collect operational status data, start and stop times, number of lifting operations, handling distance, equipment energy consumption records, and operation logs of quay cranes, yard cranes, reach stackers, forklifts, automated guided vehicles, and container trucks. S13. Collect personnel working hours records, including port entry and exit times, shift schedules, job assignments, single container working hours, and team working hours; S14. Collect financial voucher records, including expense vouchers, invoice records, energy costs, labor costs, expense allocation records, and total costs for the payment period; S15. Perform uniform format conversion processing on the collected data, standardize the timestamp, location information, and identifier fields, and output a standardized dataset.

3. The port cost tracing method based on multi-source data fusion according to claim 1, characterized in that, Step S2 includes: S21. Establish a unified time benchmark for the ship plans, berthing times, work instructions, equipment operation records, personnel working hours records and expense voucher records in the standardized dataset; S22. Based on the unified time benchmark, the granularity of the timestamps of multi-source data is adjusted, the minute-level, hour-level and day-level records are converted into a consistent time interval, and placeholder marks are filled for missing time periods, and overlapping time periods are merged in chronological order. S23. Map the berth number and berthing time according to the unified time base to generate a berth index and form a ship operation space code; S24. Map the storage yard location records according to the unified time base to generate a storage yard grid code and form a storage space code; S25. Combine the equipment number, personnel number and expense item number based on the unified time base and spatial code to generate a data index that can be used to uniquely locate the work unit; S26. Output the aligned dataset, which includes a unified time series, berth index, yard grid code, equipment number, personnel number, cost item number, and work unit index.

4. The port cost tracing method based on multi-source data fusion according to claim 1, characterized in that, Step S3 includes: S31. Extract the container number, process identifier, equipment number, personnel number, operation period, berth index and yard grid code from the aligned dataset to generate a composite identifier that can uniquely locate the operation unit. S32. Map the device energy consumption records according to the composite identifier to generate an energy consumption subset; S33. Map the personnel work hour records according to the composite identifier to generate a subset of work hours; S34. Map the financial voucher records according to the composite identifier to generate a cost subset; S35. Based on the composite identifier, the energy consumption subset, working hour subset, and cost subset are jointly bound in the time dimension, space dimension, resource dimension, and cost dimension to form a work unit dataset containing time, space, equipment, personnel, and cost information.

5. The port cost tracing method based on multi-source data fusion according to claim 1, characterized in that, S4 includes: S41. Extract equipment energy consumption records, personnel working hours records and financial voucher records from the work unit dataset, divide the equipment energy consumption status interval, working hour status interval and cost status interval, and establish an evidence set. S42. Generate a Dirichlet evidence distribution model for each type of evidence, record the parameter vectors of each state, and label the reliability weights of the evidence sources. S43. Perform prior subtraction and normalization on the parameter vectors of each evidence source, calculate the support quantity of each state, and obtain the belief value sequence and uncertainty value of each type of evidence. S44. Based on the reliability weight, the parameter vectors and belief value sequences of each evidence source are weighted and combined to generate the fused Dirichlet distribution and fused belief sequence. S45. According to the preset benchmark rate, the fused belief sequence and uncertainty are matched and synthesized, and the confidence level of each state is calculated. S46. Statistically analyze the dispersion of the belief value sequence of each evidence source for each state to obtain the inter-source dispersion sequence; calculate the deviation measure between each evidence source and the fused belief sequence to obtain the deviation sequence; synthesize the conflict degree according to the preset weight, and introduce the time overlap rate and spatial matching rate penalty coefficient for correction to generate the conflict degree of each state. S47. Output a set of evidence fusion results containing the confidence and conflict levels of each state.

6. The port cost tracing method based on multi-source data fusion according to claim 1, characterized in that, Step S5 includes: S51. Extract equipment energy consumption records, personnel work hour records, and financial voucher records from the evidence fusion results based on the composite identifier of the work unit. S52. Resample the energy consumption records of the equipment at preset time intervals, perform outlier removal and smoothing processing, generate a power sequence for the working period and form a power curve. S53. The power curve is truncated and summed according to the working period to generate energy consumption, and the peak, normal and valley periods are marked. S54. Perform cross-day segmentation, overlapping time period clipping, and shift handover marking on personnel work hour records, and aggregate work hours by work period to generate work hour quantity; S55. Perform text extraction on financial voucher records, extract expense items, amounts, voucher numbers and accounting periods, match them according to the composite identifier of the work unit, and generate expense information; S56. The energy consumption, working hours and cost information are bound to the composite identifier of the work unit, and attached to the confidence and conflict degree generated during the evidence fusion process to form a work unit feature set.

7. The port cost tracing method based on multi-source data fusion according to claim 1, characterized in that, Step S6 includes: S61. Based on the composite identifier of the work unit, gather energy consumption, working hours and cost information, and combine it with the confidence and conflict degree generated during the evidence fusion process to form a set of candidate factors; S62. Identify energy consumption factors, labor factors and cost factors in the candidate factor set respectively, and establish a cost-driving factor sequence. S63. Generate a sequential operation path based on the operation instructions, and divide the time into peak hours, normal hours and valley hours, while performing spatial segmentation in combination with the berth index and yard grid code; S64. In the improved Aumann–Shapley method, confidence is introduced as a weighting coefficient and conflict degree is introduced as a penalty factor to adjust the marginal contribution of each cost driver in different time and space segments, and generate a phased allocation value. S65. Introduce time-of-use pricing and demand charges into the calculation of phased apportionment value, and implement peak responsibility allocation and capacity apportionment for public energy consumption and costs; S66. Aggregate the phased allocation values ​​by work unit to form the allocation results of energy consumption, labor and expenses, and maintain a corresponding relationship with the total cost over the billing period; S67. Perform consistency verification and boundary truncation on the allocation results, and output the allocation value set of the job unit.

8. A port cost tracing method based on multi-source data fusion according to claim 1, characterized in that, Step S7 includes: S71. Summarize the energy consumption, working hours and allocated value based on the composite identifier of the work unit, and calculate the total cost of the work unit; S72. Combine the total cost with the berth index and the yard grid code to perform spatial mapping and generate a multi-dimensional cost distribution map of the port area. S73. Decompose the total cost according to the energy consumption factor, labor factor and expense factor, and output the factor contribution sequence; S74. Compare the factor contribution sequence with the total cost of the payment period to generate a difference test value; S75. When the difference test value exceeds the threshold, re-execute the time base correction and spatial coding calibration. S76. Re-output the updated total cost and factor contribution sequence, and generate a cost traceability result set containing cost distribution, factor contribution, and difference test values; S77. The cost traceability result set is stored in a structured manner.