Method, apparatus, device, storage medium and program for missing timing parameter inference

CN122840263APending Publication Date: 2026-09-29HANGZHOU MUHE MINGSHU TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611300270.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-26
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

对于已知基础日期和地理位置、但具体时段未记录的数据场景,其精确的时序参数取值往往无法通过单一数据源恢复,导致该参数的客观状态属于完全不确定

Benefits of technology

本申请提供的缺失时序参数推断方法,包括如下步骤:根据输入的基础输入数据生成多组候选时序参数;针对每一组候选时序参数,分别基于至少两类不同类别的原始特征数据集,通过预设的特征参数分析模型进行独立评分,得到每一组候选时序参数对应的评分值;选取评分值最高的候选时序参数作为推断结果。因此,本申请能够利用两类不同类别的原始特征数据集对输入的基础数据进行补全。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840263A_ABST
    Figure CN122840263A_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, storage medium, and program for inferring missing time-series parameters. The method includes the following steps: generating multiple sets of mutually exclusive candidate time-series parameters based on basic input data; the basic input data includes at least one of time parameters and location parameters; for each set of candidate time-series parameters, independently scoring them based on at least two different categories of original feature datasets using a preset feature parameter analysis model to obtain a score value for each set of candidate time-series parameters; determining a comprehensive ranking index and confidence level for each set of candidate time-series parameters based on the score values; and selecting the candidate time-series parameter with the highest score value and its corresponding confidence level as the inference result. Therefore, this application can complete the input basic data using two different categories of original feature datasets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of data processing technology, and in particular relates to a method for inferring missing time series parameters, a device for inferring missing time series parameters, a computer device, a computer-readable storage medium, and a computer program product. Background Technology

[0002] In data processing and information management scenarios, the precise time-series parameters for recorded time points are often completely uncertain due to incomplete data collection, missing record fields, or gaps in historical information. For data scenarios where the base date and geographical location are known, but the specific time period is not recorded, the precise time-series parameter values ​​often cannot be recovered from a single data source, resulting in a completely uncertain objective state for these parameters. This data gap problem directly affects the accuracy of subsequent data processing workflows based on complete time-series information. Existing technologies either directly reject data processing or use random assignment or default value filling methods to address data gaps. Both of these methods lead to high deviations, rendering all subsequent analytical conclusions based on incorrect parameters invalid. Therefore, how to effectively and accurately complete missing data is a technical problem that urgently needs to be solved by those skilled in the art.

[0003] The preceding description is intended to provide general background information and does not necessarily constitute prior art. Summary of the Invention

[0004] The purpose of this application is to provide a method, apparatus, computer device, computer-readable storage medium, and computer program product for inferring missing time series parameters, which can complete the input basic data using two different types of original feature datasets.

[0005] To achieve the above objectives: In a first aspect, embodiments of this application provide a method for inferring missing time-series parameters, comprising the following steps: generating multiple sets of mutually exclusive candidate time-series parameters based on input basic input data; the basic input data includes at least one of time parameters and location parameters; for each set of candidate time-series parameters, independently scoring them based on at least two different types of original feature datasets using a preset feature parameter analysis model to obtain a score value corresponding to each set of candidate time-series parameters; determining the comprehensive ranking index and confidence level of each set of candidate time-series parameters based on the score value; and selecting the candidate time-series parameter with the highest score value and its corresponding confidence level as the inference result.

[0006] In an optional embodiment of this application, multiple sets of candidate time series parameters are generated based on the input basic input data, including: parsing the basic input data to obtain existing data items; determining missing data items based on the existing data items; generating multiple candidate data items within the candidate set corresponding to the missing data items and according to a preset time granularity within a preset time period; and combining the candidate data items one by one with the existing data items to obtain multiple sets of candidate time series parameters, wherein the candidate time series parameters are different from each other.

[0007] In an optional embodiment of this application, for each group of candidate time-series parameters, independent scoring is performed based on at least two different types of original feature datasets using a preset feature parameter analysis model to obtain a score value corresponding to each group of candidate time-series parameters. This includes: processing the time-series parameters into different feature vectors based on two different types of original feature datasets; inputting the feature vectors into the feature parameter analysis model for processing to obtain a first score value and a second score value; obtaining the first score value and the second score value based on different original feature datasets; and treating the first score value and the second score value as the score value corresponding to the candidate time-series parameter.

[0008] In an optional embodiment of this application, the original feature dataset includes a first feature set and a second feature set; based on the two different categories of original feature datasets, the time series parameters are processed into different feature vectors, including: based on the first feature set, controlling the feature parameter analysis model to score the candidate time series parameters on a first time series scale to obtain a first score value; based on the second feature set, controlling the feature parameter analysis model to score the candidate time series parameters on a second time series scale to obtain a second score value; the first time series scale and the second time series scale are different.

[0009] In an optional embodiment of this application, the original feature dataset includes: a calendar time series feature set, a spatial distribution feature set, and a celestial angular distance feature set; the calendar time series feature set is used to map candidate time series parameters to feature vectors at multiple time series scales based on calendar rules and periodicity; the spatial distribution feature set is used to map candidate time series parameters to spatially related feature vectors based on geographic spatial distribution patterns; and the celestial angular distance feature set is used to map candidate time series parameters to celestial angular distance-related feature vectors based on celestial trajectories and relative positional relationships.

[0010] In an optional embodiment of this application, for each group of candidate time-series parameters, independent scoring is performed based on at least two different types of original feature datasets using a preset feature parameter analysis model to obtain a score value corresponding to each group of candidate time-series parameters. This includes: mapping a candidate time-series parameter to at least two sets of feature vectors based on the original feature dataset; inputting the feature vectors into the feature parameter analysis model for independent processing to obtain at least two sets of feature scores; and normalizing the feature scores to a preset scoring interval according to preset scoring rules to obtain a score value.

[0011] In an optional embodiment of this application, determining the comprehensive ranking index for each group of candidate time series parameters based on the score value includes: normalizing the score value to a preset score range, and weighting and summing the score values ​​corresponding to the same group of candidate time series parameters according to a preset weight to obtain the comprehensive ranking index; and / or aggregating the ranking ranks corresponding to the candidate time series parameters into a ranking aggregation index according to a preset rule, and using the ranking aggregation index as the comprehensive ranking index.

[0012] In an optional embodiment of this application, determining the comprehensive ranking index and confidence level of each group of candidate time series parameters based on the score value includes: determining two adjacent candidate time series parameters based on the comprehensive ranking index; calculating the difference between the two adjacent candidate time series parameters, and using the difference as the confidence level of the candidate time series parameter with the larger value; the difference includes at least one of difference, ratio, or ranking interval; when the candidate time series parameter with the larger value is 0, setting the confidence level of the candidate time series parameter with the larger value to 0.

[0013] In an optional embodiment of this application, the comprehensive ranking index and confidence level of each group of candidate time series parameters are determined based on the score value, including: when the confidence level corresponding to the candidate time series parameter is in the first interval, an uncertain label is attached to the candidate time series parameter, and the uncertain label will reduce the ranking of the candidate time series parameter; when the confidence level corresponding to the candidate time series parameter is in the second interval, a definite label is attached to the candidate time series parameter, and the definite label will improve the ranking of the candidate time series parameter; the value of the first interval is lower than the value of the second interval.

[0014] In an optional embodiment of this application, selecting the candidate time series parameter with the highest score and its corresponding confidence level as the inference result includes: obtaining external verification time node data; verifying the candidate time series parameter based on the external verification time node data to obtain the matching rate of each candidate time series parameter; dynamically correcting the confidence level based on the matching rate; and selecting the candidate time series parameter with the highest score and the corrected confidence level as the inference result.

[0015] In an optional embodiment of this application, the external verification time point data is independent of the feature parameter analysis model to correct the confidence level.

[0016] In an optional embodiment of this application, after selecting the candidate time series parameter with the highest score as the inference result, the process includes: incorporating the inference result into a preset historical verification sample; obtaining the consistency index of the historical verification sample; and adjusting the preset weights of the feature parameter analysis model based on the consistency index.

[0017] Secondly, embodiments of this application provide a missing time series parameter inference device, which includes: a candidate time series parameter generation module, used to generate multiple sets of candidate time series parameters based on input basic input data; a candidate time series parameter scoring module, used to independently score each set of candidate time series parameters based on at least two different types of original feature datasets using a preset feature parameter analysis model, to obtain a score value corresponding to each set of candidate time series parameters; a candidate time series parameter sorting and verification module, used to determine the comprehensive sorting index and confidence level of each set of candidate time series parameters based on the score value; and a result output module, used to select the candidate time series parameter with the highest score value and its corresponding confidence level as the inference result.

[0018] Thirdly, embodiments of this application provide a computer device, including: a processor and a memory storing a computer program, wherein when the processor runs the computer program, the steps of the above-described method are implemented.

[0019] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.

[0020] Fifthly, embodiments of this application provide a computer program product, including computer program instructions, which, when executed by a processor, implement the steps of the above-described method.

[0021] The embodiments of this application have the following beneficial effects: The missing time-series parameter inference method provided in this application includes the following steps: generating multiple sets of candidate time-series parameters based on the input basic data; for each set of candidate time-series parameters, independently scoring them based on at least two different categories of original feature datasets using a preset feature parameter analysis model to obtain a score value corresponding to each set of candidate time-series parameters; and selecting the candidate time-series parameter with the highest score value as the inference result. Therefore, this application can complete the input basic data using two different categories of original feature datasets.

[0022] The above description is merely an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it according to the contents of the specification, and to make the above and other objects, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit this application. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart illustrating a method for inferring missing timing parameters in one embodiment.

[0025] Figure 2 This is a schematic diagram of the internal functional modules of a missing timing parameter inference device provided in one embodiment.

[0026] Figure 3 This is a schematic block diagram of the structure of a computer device provided in one embodiment. Detailed Implementation

[0027] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements.

[0028] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, components, features, and elements with the same names in different embodiments of this application may have the same meaning or different meanings, the specific meaning of which must be determined by its interpretation in that specific embodiment or further in conjunction with the context of that specific embodiment.

[0029] It should be understood that although the terms first, second, third, etc., may be used herein to describe various information, this information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this document, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word “if” as used herein may be interpreted as “when…” or “in response to determination”. Furthermore, as used herein, the singular forms “a,” “an,” and “the” are intended to also include the plural forms unless the context indicates otherwise. It should be further understood that the terms “comprising,” “including,” and “including” indicate the presence of the stated feature, step, operation, element, component, item, kind, and / or group, but do not exclude the presence, occurrence, or addition of one or more other features, steps, operations, elements, components, items, kinds, and / or groups. The terms “or” and “and / or” as used herein are to be interpreted as inclusive, or mean any one or any combination thereof. Therefore, "A, B, or C" or "A, B, and / or C" means "any one of the following: A; B; C; A and B; A and C; B and C; A, B, and C". Exceptions to this definition will only occur if the combination of elements, functions, steps, or operations is inherently mutually exclusive in some way.

[0030] It should be understood that although the steps in the flowcharts of this application's embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.

[0031] It should be noted that step designations such as S110 and S120 are used in this document for the purpose of more clearly and concisely describing the corresponding content, and do not constitute a substantial limitation on the order. In specific implementation, those skilled in the art may execute S120 first and then S110, etc., but these should all be within the protection scope of this application.

[0032] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.

[0033] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.

[0034] Existing technologies cannot rely on historical data for memory-based reconstruction, nor can they utilize sensor data for multimodal joint analysis. This results in a triple challenge when attempting to complete missing time-series parameters: lack of labeled samples, lack of historical references, and lack of sensor assistance. To overcome these shortcomings, this application provides a method for inferring missing time-series parameters. For a clear description of the method provided in this embodiment, please refer to... Figures 1-3 This includes steps S110 to S140.

[0035] Step S110: Generate multiple sets of mutually exclusive candidate timing parameters based on the input basic input data; the basic input data includes at least one of time parameters and position parameters.

[0036] In one implementation, the basic input data includes at least one of a time parameter and a location parameter. Specifically, the time parameter may include, but is not limited to, a basic date field, such as year, month, and day information; the location parameter may include, but is not limited to, a geographic location field, such as longitude and latitude information. When the date format provided by the data requester is not a common format, the system can automatically convert it to a common date format. When both the time parameter and the location parameter in the basic input data are known, but the precise time series parameter (e.g., the specific hour, minute, and second) is missing, the missing time series parameter is the target parameter to be inferred.

[0037] In one embodiment, generating multiple sets of candidate time series parameters based on the input basic input data includes: parsing the basic input data to obtain existing data items; determining missing data items based on the existing data items; generating multiple candidate data items within the candidate set corresponding to the missing data items and according to a preset time granularity within a preset time period; and combining the candidate data items one by one with the existing data items to obtain multiple sets of candidate time series parameters, wherein the candidate time series parameters are different from each other.

[0038] In one implementation, the basic input data is parsed to extract existing data items, such as year, month, day, and latitude / longitude. Then, missing data items, such as time information accurate to the minute, are determined based on the existing data items. For missing data items, the system iterates through the candidate set of possible values ​​to generate multiple candidate data items. Taking a day as the inference period and minutes as the time step as an example, the system divides the entire day into 1440 candidate times (i.e., 00:00 to 23:59), with each candidate time being a candidate data item. Subsequently, the system combines each candidate data item with existing data items (date and latitude / longitude) to form a complete set of candidate time series parameters. Since each candidate data item corresponds to a different time point, the generated multiple sets of candidate time series parameters are all different from each other.

[0039] Before generating candidate time series parameters, the system determines whether time zone correction is needed based on geographic location information. Specifically, the system determines the time zone of the geographic location based on its longitude information, calculates the offset from the longitude of the standard time zone center, performs time zone correction on the date information, and generates a sequence of candidate time series parameters based on the time zone-corrected dates.

[0040] The granularity of the candidate time series parameters can be selected according to the needs of the application scenario. For example, in scenarios with sufficient computing resources and high resolution, minute-level precision can be used to generate 1440 sets of candidate time series parameters throughout the day; in scenarios with high requirements for computing efficiency, hour-level precision can be used to generate 24 sets of candidate time series parameters throughout the day; and in even finer scenarios, second-level precision can also be used. This application does not limit this.

[0041] Step S120: For each group of candidate time series parameters, independently score them based on at least two different types of original feature datasets using a preset feature parameter analysis model to obtain the score value corresponding to each group of candidate time series parameters.

[0042] In one implementation, the core of step S120 lies in using at least two different types of original feature datasets and corresponding feature parameter analysis models to independently evaluate the same set of candidate time-series parameters from different physical dimensions. The feature parameter analysis models are not interconnected and their operations are not coupled, ensuring the independence of the scoring results and the effectiveness of cross-validation.

[0043] In one embodiment, for each set of candidate time-series parameters, independent scoring is performed based on at least two different types of original feature datasets using a preset feature parameter analysis model to obtain a score value corresponding to each set of candidate time-series parameters. This includes: processing the time-series parameters into different feature vectors based on two different types of original feature datasets; inputting the feature vectors into the feature parameter analysis model for processing to obtain a first score value and a second score value; obtaining the first score value and the second score value based on different original feature datasets; and treating the first score value and the second score value as the score value corresponding to the candidate time-series parameter.

[0044] The original feature dataset includes a first feature set and a second feature set. Based on the two different categories of original feature datasets, the time series parameters are processed into different feature vectors, including: based on the first feature set, the control feature parameter analysis model scores the candidate time series parameters on the first time series scale to obtain a first score value; based on the second feature set, the control feature parameter analysis model scores the candidate time series parameters on the second time series scale to obtain a second score value; the first time series scale and the second time series scale are different.

[0045] The original feature dataset includes: a calendar time series feature set, a spatial distribution feature set, and a celestial angular distance feature set. The calendar time series feature set is used to map candidate time series parameters to feature vectors at multiple time series scales based on calendar rules and periodic patterns. The spatial distribution feature set is used to map candidate time series parameters to spatially related feature vectors based on geographic spatial distribution patterns. The celestial angular distance feature set is used to map candidate time series parameters to celestial angular distance-related feature vectors based on celestial trajectories and relative positional relationships.

[0046] In one embodiment, the independent scoring for each set of candidate time-series parameters specifically includes the following sub-steps: mapping a candidate time-series parameter to at least two sets of feature vectors based on the original feature dataset. Specifically, the system combines the candidate time-series parameter with the basic input data (date and latitude / longitude), and then substitutes them into the feature extraction rules corresponding to various types of original feature datasets to generate at least two sets of feature vectors. For example, after combining the candidate time t with the date and latitude / longitude, the system calculates the calendar time-series feature vector, the spatial distribution feature vector, and the celestial angular distance feature vector, respectively.

[0047] The feature vectors are input into the feature parameter analysis model and processed independently to obtain at least two sets of feature scores. Each model performs score calculation on the input feature vectors according to its corresponding deterministic mathematical mapping relationship and outputs the original feature scores.

[0048] The feature scores are normalized to a preset scoring range according to preset scoring rules to obtain the score value. For example, the original scores of each model are uniformly normalized to a scoring range of 0 to 100 to ensure the comparability of the score values ​​output by different models. The normalized score value is the score value of the candidate time-series parameter under the corresponding model. Furthermore, different original feature datasets will map the candidate time-series parameters to different time-series scales. The specific processing flow for each feature set, as well as the time-series scale mapped for each feature set, will be explained in detail later. The technical essence and scoring logic of the three types of feature datasets are explained below.

[0049] When scoring candidate time series parameters based on the feature parameter analysis model, the model evaluates each group of candidate time series parameters one by one and outputs a consistency score from 0 to 100.

[0050] In one embodiment, a calendar time series feature set is used to map the candidate time series parameters to feature vectors at multiple time series scales based on calendar rules and periodic patterns. The calendar time series feature set performs consistency evaluation on the candidate time series parameters based on periodic time coding fields (including the 24 solar terms boundary points, leap year markers, etc.) in a publicly available calendar dataset. The 24 solar terms coding system uses 15 degrees of solar longitude as a boundary for each solar term, and the mapping relationship between the ecliptic longitude of each solar term boundary point and the solar term code is common knowledge in the field (e.g., Winter Solstice equals 270 degrees of ecliptic longitude, equals code 0; Minor Cold equals 285 degrees of ecliptic longitude, equals code 1, and so on).

[0051] The baseline code generation method is as follows: Taking the 24 solar terms coding system as an example, firstly, the solar longitude is calculated based on the base date to determine the solar term interval to which that date belongs (e.g., the winter solstice interval is between 270 and 300 degrees). The solar longitude can be calculated using the simplified algorithm in Meeus's "Astronomical Algorithms," and the calculation accuracy is sufficient to meet the solar term boundary determination. Then, the 24-hour cycle is divided into equal-length moments with minute-level precision, and each moment corresponds to a solar term phase code. For each candidate moment, the theoretical value of its corresponding solar longitude on that date is calculated to obtain the corresponding solar term phase code. The baseline code is the solar term phase code corresponding to 0:00 on that day. Feature similarity is calculated through the one-way phase difference between the candidate code and the baseline code. The calculation method for feature similarity can be found in the following formula.

[0052] (1) In the above formula, This is the benchmark score (the highest score when the one-way phase difference is zero). This is the linear decay coefficient of the phase difference with respect to the score. This represents the unidirectional phase difference corresponding to the i-th group of candidate timing parameters. (The above...) and The value of can be set by those skilled in the art according to the width of the scoring interval and the required discrimination.

[0053] The scoring system for calendar time series feature sets can further include an auxiliary time series feature consistency score. This score is determined by the number of consistent matches in the following auxiliary coding fields: whether the candidate time is consistent with the code of the same time on the previous day, whether the candidate time is consistent with the code of the same time on the next day, and whether the candidate time is consistent with the code of the most likely typical time in the statistically significant geographical area. The aforementioned "most likely typical time in the statistically significant geographical area" is determined by the frequency of occurrence of existing time data of the corresponding geographical area and corresponding calendar node in the publicly available calendar dataset, and is a deterministic statistical result based on publicly available data.

[0054] The scoring system for the calendar time series feature set can further include a solar term boundary assignment deviation item and a time series coding structure consistency item. The solar term boundary assignment deviation item is scored according to preset rules based on the phase deviation between the candidate time and its corresponding solar term boundary point; the time series coding structure consistency item is scored according to preset rules based on the degree of conformity between the time series coding structure of the candidate time and the expected pattern. After the scores of the above items are accumulated, the score value corresponding to the calendar time series feature set is obtained after normalization, and the upper limit of the score is the maximum value of the preset score interval.

[0055] The calendar time series feature set is based on the solar term phase encoding, which mainly provides distinguishability between different hours, that is, it generates effective distinguishability at the hour-level time series scale.

[0056] A spatial distribution feature set is used to map the candidate time series parameters to spatially location-related feature vectors based on geographic spatial distribution patterns.

[0057] Specifically, the spatial distribution feature set evaluates the consistency of candidate time series parameters based on the characteristics of the solar spatial location.

[0058] The model converts the solar altitude angle into a score value according to a preset mapping function. In one exemplary implementation, for each candidate time, based on the latitude and longitude information of the base date, an astronomical algorithm calculates the spatial position angle of the sun relative to the local horizon, including the solar altitude angle and the solar azimuth angle. The solar altitude angle is the angle between the sun's direction and the local horizon, ranging from -90° to +90°. Different candidate times correspond to different continuous values ​​of the solar altitude angle. The model converts the solar altitude angle into a consistency score from 0 to 100 according to a preset mapping function, which can be found in the following formula.

[0059] (2) In the above formula, Let be the solar altitude angle corresponding to candidate time t, and sin be the sine function. This is the maximum value of the preset scoring interval (e.g., M=100). The mapping function is a deterministic mathematical expression, ensuring that different candidate times correspond to different scoring values, providing continuous discrimination at the minute-level time granularity.

[0060] It should be noted that the specific form of the above mapping function is an exemplary implementation and can be selected according to the needs of the application scenario. Its essential requirement is that the mapping is a deterministic injective function, capable of mapping continuous values ​​of solar altitude angles corresponding to different candidate times to distinct score values, thereby providing continuous discrimination at the minute-level time granularity. Those skilled in the art will understand that the final selection result of the optimal candidate time series parameters is jointly determined by the weighted fusion of all feature parameter analysis models at the result layer, rather than depending on the absolute value or monotonic direction of any single mapping function.

[0061] As an alternative implementation, the spatial distribution characteristics in the location parameters are not limited to continuous solar altitude angle values. Instead, they can be implemented using other spatially relevant features such as regional time-segment encoding based on candidate times or longitude-corrected time deviations. For example, candidate times can be encoded into discrete or continuous feature values ​​according to their day-night interval, geographical time zone segmentation, or longitude-corrected time deviation, and then converted into a consistency score of 0 to 100 according to a preset mapping rule. Both the above alternative implementations and the continuous solar altitude angle value scheme are essentially based on generating scores based on the spatial distribution characteristics corresponding to candidate times, and do not depart from the core concept of this application. Those skilled in the art can choose the appropriate spatial distribution feature implementation method according to the application scenario's requirements for time resolution and computational load. The aforementioned continuous solar altitude angle values, regional time-segment encoding, longitude-corrected time deviations, etc., are all specific implementation methods of the "spatial distribution feature set" described in this application. Those skilled in the art can choose one or combine them according to the application scenario, and all fall within the scope of the "spatial distribution feature set."

[0062] The spatial distribution feature set is based on continuous values ​​of solar altitude angle, and mainly provides discriminative power on a minute-level time scale.

[0063] A celestial angular distance feature set is used to map the candidate time series parameters to celestial angular distance-related feature vectors based on celestial body trajectories and relative positional relationships. Specifically, the consistency of candidate time series parameters is evaluated based on celestial angular distance relationships and angle parameters. The astronomical parameters used in this model (sidereal hour angle, ecliptic longitude angle, angular distance relationships, etc.) are taken from publicly available astronomical observation data. The calculation of celestial body equatorial coordinates adopts the IAU 2006 / 2000A precession nutation model (referencing the IERS Conventions 2010 standard) or VSOP87 analytical theory (Bretagnon & Francou, 1988), both of which provide a calculation accuracy of no less than 1 arcsecond, sufficient to distinguish the differences in celestial body positions between adjacent candidate times (minute-level intervals). The theoretical values ​​of celestial body trajectories are calculated based on publicly available astronomical ephemeris tables (such as the NASA JPL DE series or the NASA HORIZONS system). The model converts the lunar and solar angular distance deviations into a consistency score of 0 to 100 according to a preset mapping function, which can be found in the following formula.

[0064] (3) In the above formula, This represents the absolute deviation between the lunar and solar angular distance at candidate time t and the baseline angular distance for that day. A preset scaling factor (e.g., 11.43) is used to map the angular distance deviation to a scoring range of 0 to 100. Those skilled in the art will understand that the specific value of μ can be preset or adjusted according to the scoring discrimination requirements of the actual application scenario, and does not constitute a limitation on the scope of protection of this application. This mapping function is a deterministic mathematical expression, ensuring that different candidate times correspond to different scoring values, and has continuous discrimination at the minute-level time granularity. The reference angular distance for the day refers to the angular distance between the Sun and the Moon from a geocentric perspective at 0:00 on the day (using the same time zone reference as the candidate times). This value is uniquely calculated from publicly available astronomical ephemeris tables using a deterministic astronomical algorithm, and does not depend on human selection.

[0065] The celestial angular distance feature set is based on continuous values ​​of solar-lunar angular distance and mainly provides discriminative power on minute-level time-series scales.

[0066] It should be noted that the various feature parameter analysis models provide complementary discriminative power across different time scales. The model corresponding to the calendar time series feature set provides discriminative power based on the solar term phase encoding across different hours; the model corresponding to the spatial distribution feature set provides discriminative power based on continuous solar altitude angle values ​​at the minute level; and the model corresponding to the celestial angular distance feature set provides discriminative power based on continuous solar-lunar angular distance values ​​at the minute level. This multi-scale discriminative design ensures that at least two types of models can provide effective discriminative power when generating candidate time series parameters using different time granularities, thus guaranteeing the discriminative ability of the candidate scoring matrix.

[0067] In one implementation, each feature parameter analysis model calls different categories of original feature datasets during the scoring phase, adopts independent scoring rules, and does not share intermediate calculation results during the scoring phase.

[0068] Specifically, the aforementioned "non-sharing of intermediate calculation results" can be achieved through operational isolation between models. Each feature parameter analysis model is deployed as an independent computational unit. Each computational unit only receives shared basic input data and candidate time-series parameters through a preset data interface and returns a normalized score value to the result layer. The original feature datasets, intermediate computational variables, and internal calculation processes called by each computational unit during the scoring process are invisible to and not transmitted to other feature parameter analysis models. Thus, the scoring result of any model does not depend on the intermediate calculation processes of other models, and the scores of each model are independent of each other, effectively avoiding the contamination of the candidate scoring matrix by the same source bias.

[0069] Those skilled in the art will understand that the above-described operational isolation method is an exemplary implementation. The essence of "not sharing intermediate calculation results" is that each model does not exchange intermediate calculation data during the scoring stage and completes the scoring independently. Its specific implementation is not limited to the above-described deployment method.

[0070] In step S120, the system performs cross-validation scoring on each group of candidate time-series parameters from multiple independent physical dimensions, forming a candidate scoring matrix. The data structure of this candidate scoring matrix is ​​a two-dimensional array consisting of the score values ​​of each candidate time-series parameter under each feature parameter analysis model, and each score value is associated with the category identifier of the original feature dataset on which it is based.

[0071] It should be noted that the method involved in this application utilizes the geometric relationship of Earth's orbit (the solar term boundary is determined by the solar ecliptic longitude), the law of Earth's rotation (the physical gradient of true solar time with longitude), and the laws of celestial motion to perform deterministic calculations. The output results are candidate time series parameters with quantified confidence levels and their data consistency indices. The feature data of the above dimensions are all derived from publicly available objective physical observation data. The scoring rules within each model are deterministic calculations based on fixed mathematical expressions and do not involve subjective evaluations or behavioral suggestions.

[0072] Furthermore, the "feature parameter analysis model" referred to in this application refers to a computational unit that performs deterministic scoring calculations on the input feature vector based on a fixed mathematical expression. Its scoring rules are determined by a pre-defined mathematical mapping relationship and do not include or depend on machine learning models or neural network models learned from training samples. The intermediate calculation results of each feature parameter analysis model during the scoring stage are not visible to or propagated to other models to ensure the independence of the scoring results of each model and to avoid contamination of the relative ranking of the candidate scoring matrix by homology bias.

[0073] It should also be noted that the "basic input data" described in this application includes at least a basic date field and a geographic location field. However, in specific applications, other redundant fields or additional information that do not affect the core inference logic of this application may also be received. Whether or not such additional information is received does not change the essential process of deterministic calculation based on the basic date and geographic location of each feature parameter analysis model, nor does it change the essential utilization of the technical solution of this application. The different categories of original feature datasets involved in this application are not limited to the above-mentioned calendar time series feature set, spatial distribution feature set, and celestial angular distance feature set. Those skilled in the art can select any publicly available, objective dataset with deterministic calculation rules and whose values ​​change deterministically with the changes of candidate time series parameters as the original feature data source, according to the actual application scenario. Based on this, the core concept of this application—candidate time generation, multi-source independent scoring, result layer fusion, and confidence quantification—can be extended to scenarios where other publicly available time-related data are used as feature sources. Those skilled in the art understand that there is a deterministic calculation relationship between the above-mentioned data sources and candidate times, which can replace or supplement the specific feature data sources in the foregoing embodiments of this application without departing from the core concept of this application. Step S130: Determine the comprehensive ranking index and confidence level of each group of candidate time series parameters based on the score value.

[0074] Step S140: Select the candidate time series parameter with the highest score and the corresponding confidence level as the inference result.

[0075] In one embodiment, determining the comprehensive ranking index for each group of candidate time series parameters based on the score value includes: normalizing the score value to a preset score range, and weighting and summing the score values ​​corresponding to the same group of candidate time series parameters according to a preset weight to obtain the comprehensive ranking index; and / or, aggregating the ranking ranks corresponding to the candidate time series parameters into a ranking aggregation index according to a preset rule, and using the ranking aggregation index as the comprehensive ranking index.

[0076] In one embodiment, step S130 includes: weighting and summing the score values ​​according to the preset weights corresponding to each of the original feature datasets to obtain a comprehensive ranking index. Specifically, the calculation formula for the comprehensive ranking index is shown in the following formula.

[0077] (4) In the above formula, The preset weights for the feature parameter analysis model corresponding to the i-th set of original feature datasets are: Let be the normalized score of the i-th model for the candidate time series parameter t. In a preferred embodiment of this application, when three core feature parameter analysis models are set, the weight of the first feature parameter analysis model (calendar time series dimension) is configured as 50%, the weight of the second feature parameter analysis model (spatial distribution dimension) is configured as 30%, and the weight of the third feature parameter analysis model (celestial angular distance dimension) is configured as 20%. The above weights are exemplary preferred configurations and can be adjusted according to the consistency index of historical verification samples based on the actual application scenario.

[0078] In one embodiment, determining the comprehensive ranking index and confidence level of each group of candidate time series parameters based on the score value includes: determining two adjacent candidate time series parameters based on the comprehensive ranking index; calculating the difference between the two adjacent candidate time series parameters, and using the difference as the confidence level of the candidate time series parameter with the larger value; the difference includes at least one of difference, ratio, or ranking interval; when the candidate time series parameter with the larger value is 0, setting the confidence level of the candidate time series parameter with the larger value to 0.

[0079] In one implementation, specifically, the system sorts all candidate time-series parameters according to a comprehensive ranking index, determines two adjacent candidate time-series parameters (i.e., the two candidates that are immediately next to each other after ranking), and calculates the confidence score based on the difference between the larger and second-largest values ​​of the two adjacent candidate time-series parameters. The formula for calculating the confidence score can be found in the following equation.

[0080] (5) In the above formula, For confidence level, As the highest comprehensive ranking indicator, It is the second highest comprehensive ranking indicator. When the value equals 0, the confidence score C is assigned a value of 0. The physical meaning of this confidence score calculation method is that when multiple independent models give highly consistent scores to the same candidate time-series parameter, and the candidate's overall ranking index far exceeds that of the second-best candidate, it indicates that the independent scoring dimensions have formed strong cross-validation for the candidate, and the inference result is highly reliable. Conversely, if the overall ranking indices of the best and second-best candidates are close, it indicates that the models have significant disagreements, and the inference result is less reliable.

[0081] It should be noted that the confidence level mentioned above is based on the score discrimination between candidates and is applicable to scenarios where the true value of the time series parameter to be inferred is completely missing and there is no labeled true value available for comparison. It differs from the confidence level calculation method based on labeled sample prediction probability calibration in terms of technical approach: the latter takes the consistency between the prediction result and the true label as the optimization goal, while the former takes the degree of distinguishability between the best candidate and the second-best candidate as the measurement goal. The problems they solve and the technical means they use are fundamentally different.

[0082] In one embodiment, the comprehensive ranking index and confidence level of each group of candidate time series parameters are determined based on the score value, including: when the confidence level corresponding to the candidate time series parameter is in the first interval, an uncertain label is attached to the candidate time series parameter, and the uncertain label will reduce the ranking of the candidate time series parameter; when the confidence level corresponding to the candidate time series parameter is in the second interval, a definite label is attached to the candidate time series parameter, and the definite label will improve the ranking of the candidate time series parameter; the value of the first interval is lower than the value of the second interval.

[0083] In one embodiment, the inference results are graded according to the preset numerical range in which the quantified confidence level falls. Specifically: when the confidence level corresponding to the candidate time series parameter is in the first range (i.e., the lower confidence range), an uncertainty label is attached to the candidate time series parameter. This uncertainty label lowers the ranking of the candidate time series parameter, indicating to the subsequent system that the inference result has high uncertainty. When the confidence level corresponding to the candidate time series parameter is in the second range (i.e., the higher confidence range), a certainty label is attached to the candidate time series parameter. This certainty label improves the ranking of the candidate time series parameter, indicating that the subsequent system can use this result as a high-confidence candidate result in the subsequent data application process. The values ​​in the first range are lower than those in the second range.

[0084] In one embodiment, selecting the candidate time series parameter with the highest score and its corresponding confidence level as the inference result includes: acquiring external verification time node data; verifying the candidate time series parameter based on the external verification time node data to obtain the matching rate of each candidate time series parameter; dynamically correcting the confidence level based on the matching rate; and selecting the candidate time series parameter with the highest score and the corrected confidence level as the inference result.

[0085] In one implementation, external validation time point data is independent of the feature parameter analysis model to adjust the confidence level.

[0086] External verification time node data can specifically be important, known event records with clear time stamps (such as equipment status change time nodes, business process node change time nodes, data record version change time nodes, entity registration time nodes, voucher issuance time nodes, status transition record time nodes, etc.). The system uses these events to verify the inference results—checking whether the inferred time sequence parameters match these events. A single event may match individually (many time points can match the same event), but the probability of matching a dozen or so events is extremely low. The external verification time node data and the base date may belong to different dates. The system aligns it to the same time coordinate system as the candidate time sequence parameters according to a preset periodic mapping rule for consistency verification.

[0087] It is worth noting that the confidence level correction based on the external validation time-node data is independent of the feature parameter analysis model. In other words, the correction operation on the external validation time-node data, together with the feature parameter analysis model, constitutes a multi-level time-series validation architecture. Furthermore, external event validation does not participate in the fusion calculation of the feature parameter analysis model's result layer.

[0088] Specifically, the matching rate correction confidence process involves first mapping candidate time-series parameters to a time-series coding coordinate system to obtain the corresponding time node sequence; then mapping event timestamps from external verification time node data to the same coordinate system, and calculating the time-series distance (unit consistent with the period scale) between each event timestamp and the nearest time node. When the external verification event timestamp differs from the base date, alignment is performed according to the period mapping rule: the hour-minute-second portion of the event timestamp is mapped to the corresponding moment in the 24-hour coordinate system of the base date, i.e., with a period length of 24 hours, the remainder after modulo 24 hours is used as its mapping moment in the base date coordinate system.

[0089] The following formula can be used to calculate the matching rate (consistency matching value).

[0090] (6) In the above formula, Let be the temporal distance between the timestamp of the j-th event and the nearest time node. This is the preset tolerance threshold.

[0091] The matching rate E is calculated as follows.

[0092] (7) in Let j be the event weight parameter for the j-th external verification time node. The reliability coefficient of the event source. This represents the consistency match value between the j-th event and the candidate time series parameters (range 0 to 1). When the event weight parameter... If not provided, the default value is 1; when the event source reliability coefficient is... If not provided, the default value is 1.

[0093] The system uses the matching rate E to correct the inference results of the candidate time series parameters: if all external verification events are highly matched with the recommended candidate results, the matching rate is high and the confidence is further improved; if there are mismatched events, the matching rate is low and the system lowers the confidence index of this inference.

[0094] In one embodiment, the calculation method for the corrected confidence level can refer to the following formula.

[0095] (8) In the above formula, This represents the baseline confidence level obtained earlier based on the rating discrimination. To calibrate the intensity coefficient, The above two parameters are preset values, serving as a consistency threshold. In a preferred embodiment of this application, the calibration intensity coefficient... The value range is [5, 20], and the consistency threshold is... The value range is [0.4, 0.8]. The larger the value, the greater the correction of the confidence level by the external validation; The larger the value, the higher the threshold for external verification. The above value range is an exemplary preferred configuration, and those skilled in the art can select the appropriate value based on the actual application scenario and the required verification sensitivity.

[0096] In one embodiment, after selecting the candidate time series parameter with the highest score as the inference result, the process includes: incorporating the inference result into a preset historical verification sample; obtaining the consistency index of the historical verification sample; and updating the feature parameter analysis model based on the consistency index.

[0097] In one implementation, the consistency index refers to the average degree of agreement between the scoring results of each feature parameter analysis model and the known correct time-series parameters in historical validation samples, or the average consistency score E between the scoring results of each model and the corresponding known actual time-series parameter validation events. A higher consistency index value indicates that the evaluation results of the corresponding model on historical samples are more reliable, and the preset weight of the model is increased accordingly. By iteratively and dynamically adjusting the fusion weights of each model based on the consistency index of historical validation samples, the system can optimize the fusion weight configuration of each model according to historical validation results during long-term operation.

[0098] It should be noted that the external validation is used to correct the confidence index of the current inference and does not constitute a cross-inference adjustment mechanism for the model fusion weights. The adjustment of the model fusion weights is driven by the consistency index of historical validation samples, and the two are independent of each other in terms of time scale and functional positioning.

[0099] Based on the same inventive concept as the foregoing embodiments, the following detailed description of the foregoing embodiments is provided through a specific example. This example uses a data record lacking specific time-series parameters as an example to fully demonstrate the calculation process from input data to output confidence. It should be noted that all numerical values ​​involved in the embodiments of this specification are deterministic calculation demonstration results based on publicly available calendar data and astronomical data, used to illustrate the feasibility of each scoring rule and fusion process; these numerical values ​​are not based on actual experimental results of labeled samples, nor do they constitute any claim to the inference accuracy or inference performance of this application.

[0100] Input data: Base date: January 1, 2000, located in a certain city, with a geographical coordinate of xx degrees north latitude and xx degrees east longitude.

[0101] Step S110: Generate candidate time-series parameters for the entire day with minute-level precision. This numerical example demonstrates the calculation process and confidence formula using candidate times of 06:00 and 14:00 as examples. Complete screening requires calculating all 1440 candidate times one by one. The optimal candidate is determined by the highest comprehensive ranking index after the joint fusion of the three model results layers. Candidates outside the scope of this example are declared as globally optimal. The calculation logic for the remaining candidates is the same.

[0102] Step S120 – Calendar time series feature set scoring.

[0103] According to astronomical algorithms, the sun's ecliptic longitude on January 1, 2000, was 280 degrees, falling within the winter solstice range (270 to 300 degrees). The phase coding for the 24 solar terms uses 0 as the starting point for the winter solstice. The one-way phase difference between the candidate time 06:00 and the winter solstice starting point is 6 hours; the one-way phase difference between the candidate time 14:00 and the winter solstice starting point is 14 hours.

[0104] Baseline code (corresponding to 0:00 on the current day) = 0. Feature similarity. ,in =0.5. Based on the above calculations, ; In the above formula, the baseline value of 30 is the scoring benchmark for feature similarity scores, corresponding to the highest score when the one-way phase difference is zero; the coefficient... The linear attenuation coefficient of the phase difference with respect to the score is taken in this embodiment. =0.5.

[0105] The auxiliary temporal feature consistency score is determined by the number of consistent matches in the following three auxiliary coding fields: 5 points if the candidate time matches the code of the same time on the previous day; 5 points if it matches the code of the same time on the next day; and 10 points if it matches the code of the most statistically likely typical time in the geographical area. The sum of these three scores is the auxiliary temporal feature consistency score, with a maximum score of 20 points. The auxiliary temporal feature consistency scores corresponding to the two candidate times are 13 and 8, respectively.

[0106] The scoring system for the calendar time series feature set includes a solar term boundary assignment deviation item (score cap of 25 points) and a time series coding structure consistency item (score cap of 25 points). For the example times selected in this embodiment, 0:00 on the day and the candidate times 06:00 and 14:00 are all within the same solar term interval. The solar term boundary assignment deviation is negligible within the tolerance range, and the time series coding structure consistency does not produce additional distinguishing effect between the selected example times at the hour-level precision. Therefore, the above two items assign 0 points to both sets of candidate times.

[0107] The sum of the four factors—feature similarity (27.0 and 23.0), auxiliary temporal feature coordination (13 and 8), solar term boundary attribution deviation (0 and 0), and temporal coding structure consistency (0 and 0)—results in a score of 40.0 and 31.0 for the calendar temporal feature set (with a maximum score of 100).

[0108] Step S120 – Spatial distribution feature set scoring.

[0109] According to astronomical algorithms, the sunrise time for the city that day was 07:36 and the sunset time was 16:53. For the candidate time 06:00, the sun was below the horizon with a solar altitude angle of -17.8 degrees; for the candidate time 14:00, the sun was above the horizon with a solar altitude angle of 22.7 degrees. Based on the mapping function... calculate: ; Different candidate times correspond to different solar spatial locations, thus the score values ​​have continuous discriminative power at the minute level. The score values ​​corresponding to the spatial distribution feature set are 65.3 and 30.7 (score cap 100).

[0110] Step S120 – Scoring of celestial angular distance feature set.

[0111] The theoretical value of the solar-moon angular distance on January 1, 2000, varies with time. Astronomical algorithms calculated that the deviation between the solar-moon angular distance at candidate time 06:00 and the baseline angular distance for that day was 2.8 degrees, and the deviation at candidate time 14:00 was 6.4 degrees. According to the mapping function... ,Pick =11.43, calculated as follows: ; The scores for the celestial angular distance feature set are 68.0 and 26.8 (score cap 100).

[0112] Step S130 – Weighted fusion.

[0113] According to preset weights ( =0.5, =0.3, =0.2) Perform weighted summation fusion. In the two candidate time points selected in this example: (9) (10) That is, the comprehensive ranking index for candidate time 06:00 is 53.19, and the comprehensive ranking index for candidate time 14:00 is 30.07.

[0114] To demonstrate the confidence formula, the higher of the two values ​​is taken as the highest comprehensive ranking index. The lower value is used as the second highest comprehensive ranking indicator. Substitute into subsequent calculations. During complete filtering. and The highest and second-highest values ​​among all candidate times should be taken, which are not within the scope of this example demonstration.

[0115] Step S130 – Confidence calculation.

[0116] (11) Step S140 – External verification.

[0117] Consistency verification was performed using an external verification event (January 1, 2000, 06:30). The time-series distance was calculated using the following formula.

[0118] (12) further, Take 2 hours, and refer to the following formula for the matching rate process.

[0119] (13) Set event weight parameters Reliability coefficient of event source The external validation consistency score is obtained based on the calculated matching rate. Calibration intensity coefficient Consistency threshold The confidence level correction calculation process can be referred to the following formula.

[0120] =45.0% (14) The numerical examples above fully demonstrate the entire computational process from basic input data to candidate scoring matrix generation, weighted fusion, confidence calculation, and external verification correction. In the validation using real historical data, historical data records with known actual time-series parameters were used as the validation benchmark. A consistency comparison was performed between the optimal candidate time-series parameters inferred by the system and the actual time-series parameters, and a consistency index was calculated. This validation process shows that the method of this application can provide candidate completion results with quantifiable confidence for missing time-series parameters under conditions of no labeled samples and no historical references. Its consistency index can be used to evaluate the reliability of the inference results. It should be noted that the above consistency index measures the degree of matching between the system output results and known actual time-series parameters, which is different from the prediction accuracy index based on labeled sample training, and does not constitute any quantitative claim on the inference performance of this application.

[0121] Therefore, existing ensemble learning schemes typically perform multi-model scoring and fusion for a fixed feature space, with each model's weights using a fixed configuration or adaptively adjusted based on the model's internal error signals. The substantial difference between this scheme and the aforementioned existing technologies in solving the problem of missing time-series parameter inference lies in the following: First, under the constraint of only two input fields—basic date and geographic location—this scheme selects three types of data sources with deterministic computational relationships to date and geographic location: calendar time-series feature sets, spatial distribution feature sets, and celestial angular distance feature sets. It also designs a multi-scale complementary scoring architecture, where the calendar model provides hourly-level discrimination, while the spatial distribution model and celestial angular distance model provide minute-level discrimination, ensuring that the candidate scoring matrix possesses effective discriminative capabilities at different time granularities. This selection and multi-scale complementary architecture design are not standard practices in existing missing value imputation or ensemble learning schemes. In existing technologies, astronomical calendar data is used for forward calculation (calculating celestial positions from known times), but not for backward inference (inferring missing times from known positions); ensemble learning is used for supervised classification and regression tasks, but not for unsupervised physical inference scenarios; multiple imputation is used for statistical missing value imputation and does not involve physical deterministic constraints. The aforementioned differences make the application scenarios and technical paths of this application fundamentally different from those of existing technologies.

[0122] Second, this scheme verifies the temporal consistency of the optimal candidate by receiving external verification time node data, and dynamically adjusts the confidence index of the current inference based on the verification results. The preset weights of each feature parameter analysis model during result layer fusion can be iteratively and dynamically adjusted according to the consistency index of historical verification samples, so that the weights of each model can be optimized based on historical verification results.

[0123] Third, the weights, scaling factors, and other parameters in this scheme are merely illustrative parameters used to demonstrate the calculation process, and their specific values ​​do not constitute a limitation on the scope of protection of this application. Those skilled in the art can select, preset, or adjust these parameters within the formulas and recommended values ​​given in the specification, according to the actual application scenario.

[0124] The above comparative analysis is only used to help understand the differences between this application and the prior art, and does not constitute a limitation on the scope of protection of this application.

[0125] In summary, the method provided in the above embodiments cross-validates candidate time-series parameters from different physical dimensions using multiple independent models, enabling the inference results to obtain cross-verification from different types of data sources and reducing dependence on a single data source. Quantifying the confidence level output improves the interpretability of the output results, facilitating subsequent hierarchical processing based on the confidence level. Introducing external verification time node data for time-series consistency verification further enhances the consistency between the inference results and external objective events. Iteratively and dynamically adjusting the fusion weights of each model based on consistency indicators from historical verification samples allows the system to optimize the fusion weight configuration of each model based on historical verification results during long-term operation. The entire method can be deployed on the server side for automated execution, possessing scalability and stability suitable for industrial applications.

[0126] Therefore, the method provided in this application can effectively infer and complete missing time series parameters and output inference results with quantified confidence, solving the technical problem that the prior art cannot effectively infer key time series parameters in scenarios where they are completely missing.

[0127] In one embodiment, the technical solution of this application can be further extended to support an architecture that supports more than three sets of feature parameter analysis models. In the extended architecture, each newly added model also follows the principle of independent scoring, that is, it calls its own original feature dataset, uses independent feature calculation rules, and does not share intermediate calculation results with other models. All models are only weighted and fused at the result layer. The core reasoning logic of this extended architecture is completely consistent with the three-model architecture. The increase in the number of models does not change the inventive concept of this application and is a technical extension that can be naturally implemented by those skilled in the art based on the disclosure of this application.

[0128] Those skilled in the art will understand that the aforementioned different categories of original feature datasets are not limited to the calendar time series feature sets, spatial distribution feature sets, and celestial angular distance feature sets. Those skilled in the art can select any publicly available, objective dataset with defined calculation rules and whose values ​​exhibit deterministic changes with the candidate time series parameters as the original feature data source, based on the actual application scenario. On this basis, the core concepts of this application—candidate time generation, multi-source independent scoring, result-level fusion, and confidence quantification—can be extended to scenarios where other publicly available time-related data serve as feature sources.

[0129] The quantitative confidence level described in this application is based on the score discrimination between candidates and is applicable to scenarios where the true value of the time series parameter to be inferred is completely missing and there is no labeled true value available for comparison. It differs from the confidence level calculation method based on the prediction probability calibration of labeled samples in terms of technical approach: the latter takes the consistency between the prediction result and the true label as the optimization goal, while the former takes the degree of distinguishability between the best candidate and the second-best candidate as the measurement goal. The problems they solve and the technical means they use are fundamentally different.

[0130] This solution provides candidate completion results with quantified confidence for scenarios where key time-series parameters are missing, and achieves the following beneficial effects: (1) Improved inference stability: By cross-validating at least two models based on original feature datasets of different categories, the impact of homology bias on the output results can be reduced compared to a single model; (2) Quantified confidence output: The introduction of quantified confidence indicators from 0 to 100% improves the interpretability of the output results, making it easier for the subsequent system to perform hierarchical processing based on confidence; (3) Dynamic calibration capability: By verifying the temporal consistency of external verification data, the consistency between the inference results and external objective events is improved; (4) Industrial deployment capability: The entire method can be deployed in an industrial environment. Deployed on the server side for automated execution, it has high scalability and stability; (5) It adopts a model scoring system based on different categories of original feature datasets. Each model independently obtains features from public datasets in different fields. During the scoring stage, intermediate calculation results are not shared, effectively avoiding homogeneity bias; (6) The candidate discrimination confidence is calculated by using the difference between the highest score and the second highest score, which directly reflects the degree of dispersion between the best candidate and the second best candidate, and is more suitable for multi-candidate time series parameter inference scenarios; (7) The fusion weights of each model are dynamically adjusted iteratively based on the consistency index of historical verification samples. In the long-term operation, the system can optimize the fusion weight configuration of each model according to the historical verification results.

[0131] Furthermore, this application also provides a missing timing parameter inference device, which specifically includes the following functional modules. For a clear description of the module architecture of the missing timing parameter inference device 20 provided by the method of this application, please refer to... Figure 2 As shown.

[0132] The candidate timing parameter generation module 21 is used to generate multiple sets of candidate timing parameters based on the input basic input data.

[0133] The candidate time series parameter scoring module 22 is used to independently score each group of candidate time series parameters based on at least two different types of original feature datasets using a preset feature parameter analysis model, so as to obtain the score value corresponding to each group of candidate time series parameters.

[0134] The candidate time series parameter sorting and verification module 23 determines the comprehensive sorting index and confidence level of each group of candidate time series parameters based on the score value.

[0135] The result output module 24 is used to select the candidate time series parameter with the highest score and the corresponding confidence level as the inference result.

[0136] In one embodiment, the device further includes a consistency verification module, used to acquire external verification time node data, perform time series consistency verification on candidate time series parameters, and dynamically correct the confidence index.

[0137] Figure 3 An internal structural diagram of a computer device in one embodiment is shown. This computer device can specifically be a terminal or a server. Figure 3 As shown, the device includes: a processor 310 and a memory 311 storing a computer program; wherein, Figure 3 The processor 310 shown in the diagram does not indicate that there is only one processor 310, but only indicates the positional relationship of the processor 310 relative to other devices. In practical applications, there can be one or more processors 310; similarly, Figure 3 The memory 311 illustrated herein has the same meaning, that is, it is only used to indicate the positional relationship of memory 311 relative to other devices. In practical applications, there can be one or more memories 311. When the processor 310 runs the computer program, the method applied to the above-mentioned device is implemented.

[0138] The device may also include at least one network interface 312. The various components of the device are coupled together via a bus system 313. It is understood that the bus system 313 is used to implement communication between these components. In addition to a data bus, the bus system 313 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 3 The general designated all buses as Bus System 313.

[0139] The memory 311 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memory 311 described in the embodiments of this application is intended to include, but is not limited to, these and any other suitable types of memory.

[0140] The memory 311 in this embodiment is used to store various types of data to support the operation of the device. Examples of this data include: any computer programs used to operate on the device, such as operating systems and applications; contact data; phonebook data; messages; pictures; videos, etc. The operating system includes various system programs, such as the framework layer, core library layer, driver layer, etc., used to implement various basic services and handle hardware-based tasks. Applications can include various applications, such as media players, browsers, etc., used to implement various application services. Here, the program implementing the method of this embodiment can be included in the application.

[0141] Based on the same inventive concept as the foregoing embodiments, this embodiment also provides a computer-readable storage medium storing a computer program. The computer-readable storage medium can be a magnetic random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM), etc.; it can also be various devices including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc. When the computer program stored in the computer-readable storage medium is run by a processor, it implements the above method. For the specific steps implemented when the computer program is executed by the processor, please refer to [link to relevant documentation]. Figure 1 The description of the illustrated embodiments will not be repeated here.

[0142] This application also provides a computer program product, including computer program instructions, which, when executed by a processor, implement the steps of the above-described method.

[0143] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0144] In this document, the terms “including,” “comprising,” or any other variations thereof are intended to cover non-exclusive inclusion, which includes not only the elements listed but also other elements not expressly listed.

[0145] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for inferring missing time series parameters, characterized in that, Includes the following steps: Multiple sets of mutually exclusive candidate timing parameters are generated based on the input basic input data; the basic input data includes at least one of time parameters and position parameters; For each set of candidate time series parameters, based on at least two different types of original feature datasets, a preset feature parameter analysis model is used to independently score them, thereby obtaining the score value corresponding to each set of candidate time series parameters. Based on the score, a comprehensive ranking index and confidence level for each group of candidate time series parameters are determined; The candidate time series parameter with the highest score and its corresponding confidence level are selected as the inference result.

2. The method for inferring missing time series parameters as described in claim 1, characterized in that, The process of generating multiple sets of mutually exclusive candidate timing parameters based on the input basic data includes: Parse the basic input data to obtain existing data items; determine missing data items based on the existing data items; Within the candidate set corresponding to the missing data item, multiple candidate data items are generated according to a preset time granularity within a preset time period; The candidate data items are combined one by one with the existing data items to obtain multiple sets of candidate time series parameters, and the candidate time series parameters are different from each other.

3. The method for inferring missing time series parameters as described in claim 1, characterized in that, For each group of candidate time-series parameters, based on at least two different categories of original feature datasets, independent scoring is performed using a pre-defined feature parameter analysis model to obtain a score value corresponding to each group of candidate time-series parameters, including: Based on the two different categories of the original feature datasets, the time-series parameters are processed into different feature vectors; The feature vector is input into the feature parameter analysis model for processing to obtain a first score value and a second score value; the first score value and the second score value are obtained based on different original feature datasets. The first score and the second score are regarded as the score values ​​corresponding to the candidate time series parameters.

4. The method for inferring missing time series parameters as described in claim 3, characterized in that, The original feature dataset includes a first feature set and a second feature set; Based on the original feature datasets of two different categories, the time-series parameters are processed into different feature vectors, including: Based on the first feature set, the feature parameter analysis model is controlled to score the candidate time series parameters on the first time series scale to obtain a first score value; Based on the second feature set, the feature parameter analysis model is controlled to score the candidate time series parameters on the second time series scale to obtain a second score value; The first time scale and the second time scale are different.

5. The method for inferring missing time series parameters as described in claim 1, characterized in that, The original feature dataset includes: calendar time series feature set, spatial distribution feature set, and celestial angular distance feature set; The calendar time series feature set is used to map the candidate time series parameters to feature vectors at multiple time series scales based on calendar rules and periodic patterns. The spatial distribution feature set is used to map the candidate time series parameters to spatially location-related feature vectors based on geographic spatial distribution patterns. The celestial angular distance feature set is used to map the candidate time series parameters to celestial angular distance-related feature vectors based on the celestial trajectory and relative position relationship.

6. The method for inferring missing time series parameters as described in claim 1, characterized in that, For each group of candidate time-series parameters, based on at least two different categories of original feature datasets, independent scoring is performed using a pre-defined feature parameter analysis model to obtain a score value corresponding to each group of candidate time-series parameters, including: Based on the original feature dataset, a candidate time-series parameter is mapped to at least two sets of feature vectors; The feature vectors are input into the feature parameter analysis model and processed independently to obtain at least two sets of feature scores; The feature scores are normalized to a preset scoring range according to preset scoring rules to obtain the score value.

7. The method for inferring missing time series parameters as described in claim 1, characterized in that, The step of determining the comprehensive ranking index for each group of candidate time-series parameters based on the score values ​​includes: The score values ​​are normalized to a preset score range, and the score values ​​corresponding to the same group of candidate time-series parameters are weighted and summed according to preset weights to obtain the comprehensive ranking index; and / or, The rankings corresponding to the candidate time series parameters are aggregated into a ranking aggregation index according to a preset rule, and the ranking aggregation index is used as the comprehensive ranking index.

8. The method for inferring missing time series parameters as described in claim 1, characterized in that, The determination of the comprehensive ranking index and confidence level of each group of candidate time series parameters based on the score value includes: Based on the comprehensive ranking index, determine two adjacent candidate time series parameters; Calculate the difference between two adjacent candidate time series parameters, and use the difference as the confidence level of the candidate time series parameter with the larger value; the difference includes at least one of difference, ratio, or rank interval; When the candidate time series parameter with the larger value is 0, the confidence level of the candidate time series parameter with the larger value is set to 0.

9. The method for inferring missing time series parameters as described in claim 1, characterized in that, The determination of the comprehensive ranking index and confidence level of each group of candidate time series parameters based on the score value includes: When the confidence level corresponding to the candidate time series parameter is in the first interval, an uncertainty label is attached to the candidate time series parameter, and the uncertainty label will reduce the ranking of the candidate time series parameter. When the confidence level corresponding to the candidate time series parameter is in the second interval, a determination label is attached to the candidate time series parameter, and the determination label will improve the ranking of the candidate time series parameter; the value of the first interval is lower than the value of the second interval.

10. The method for inferring missing time series parameters as described in claim 1, characterized in that, The step of selecting the candidate time-series parameter with the highest score and the corresponding confidence level as the inference result includes: Obtain external verification time node data; The candidate time series parameters are verified based on the external verification time node data to obtain the matching rate of each candidate time series parameter; The confidence level is dynamically adjusted based on the matching rate. The candidate time series parameter with the highest score and the corrected confidence level are selected as the inference result.

11. The method for inferring missing time series parameters as described in claim 10, characterized in that, The external verification time point data is independent of the feature parameter analysis model, in order to correct the confidence level.

12. The method for inferring missing time series parameters as described in claim 1, characterized in that, After selecting the candidate time-series parameter with the highest score as the inference result, the process includes: The inference results are incorporated into a preset historical verification sample; Obtain the consistency index of the historical verification samples; The preset weights of the feature parameter analysis model are adjusted based on the consistency index.

13. A device for inferring missing timing parameters, characterized in that, The missing time series parameter inference includes: The candidate time series parameter generation module is used to generate multiple sets of candidate time series parameters based on the input basic data. The candidate time series parameter scoring module is used to independently score each group of candidate time series parameters based on at least two different types of original feature datasets using a preset feature parameter analysis model, so as to obtain the score value corresponding to each group of candidate time series parameters. The candidate time series parameter sorting and verification module determines the comprehensive sorting index and confidence level of each group of candidate time series parameters based on the score value; The result output module is used to select the candidate time series parameter with the highest score and the corresponding confidence level as the inference result.

14. A computer device, characterized in that, Including processor and memory; The processor is configured to execute a computer program stored in the memory to implement the method as described in any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 12.

16. A computer program product comprising computer program instructions, characterized in that, When the computer program instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 12.