Second-hand car data set construction method and device based on data management and storage medium
By cleaning, standardizing, and performing bibasic vector decomposition on used car transaction data, and combining it with a physical constraint model, the problem of difficulty in preserving the correlation between operational behavior and state changes in used car data was solved, generating a structured dataset that conforms to the vehicle usage logic and improving data utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUIZHOU ENGINE TECHNOLOGY IND CO LTD
- Filing Date
- 2025-12-25
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies struggle to fully preserve the correlation between operational behaviors and status changes in used car text data, and lack a systematic check on whether data records conform to objective laws during vehicle use, resulting in low data utilization.
By cleaning and standardizing the original vehicle description text in the used car transaction scenario, extracting event units and performing bibasic vector decomposition, and combining the geometric constraint model of the vehicle's physical state for detection and correction, a structured data unit aligned with the used car valuation model is finally generated.
It achieves precise capture of vehicle operation behavior and state changes, breaks through the limitations of causal correlation, ensures that the data conforms to the objective physical logic in the process of vehicle use, and improves the semantic integrity, quantitative accuracy and business adaptability of the dataset.
Smart Images

Figure CN121996774A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method, apparatus and storage medium for constructing a used car dataset based on data governance. Background Technology
[0002] With the advancement of digitalization in the used car market, building standardized used car datasets based on unstructured text data has become a fundamental technology for supporting vehicle valuation and transaction risk control. The core requirement lies in extracting effective information reflecting changes in vehicle condition from text data such as vehicle usage records and maintenance descriptions. Currently, related technologies for processing used car text data often extract discrete attributes through word segmentation, entity recognition, and keyword matching, then fill the extracted attribute values into a preset data template to obtain a structured dataset containing basic parameters. However, when faced with dynamic descriptions of vehicle usage, this approach often fails to fully preserve the correlation between operational behaviors and state changes, resulting in the inability to effectively transform the implicit causal patterns in the data into usable features. Furthermore, existing technologies lack a systematic check to ensure that data records conform to objective laws during vehicle use, potentially leading to abnormal data that does not align with the vehicle's physical characteristics or usage logic. These issues limit the current used car datasets in supporting accurate analysis and decision-making, making it difficult to fully realize the potential value of unstructured text data. Summary of the Invention
[0003] In view of this, the present invention provides a method, apparatus, and storage medium for constructing a used car dataset based on data governance. The technical solution of the embodiments of the present invention is implemented as follows: On one hand, embodiments of the present invention provide a method for constructing a used car dataset based on data governance. The method includes: acquiring original vehicle description text from a used car transaction scenario; performing text cleaning and standardization on the original vehicle description text to obtain a set of vehicle description texts to be processed, the set of vehicle description texts to be processed containing state description information and operation record information of the vehicle at different stages of use; extracting event units from the set of vehicle description texts to be processed, identifying and extracting event units containing operation behaviors and state changes in the set of vehicle description texts to be processed, the event units being used to characterize the state change process of the vehicle at the corresponding time node; and performing bibasic vector decomposition on each event unit to decompose the event unit. The system consists of a bibasic vector composed of an operation vector and a state vector. The operation vector represents the force and direction of the active intervention applied to the vehicle, while the state vector represents the magnitude of the vehicle attribute shift caused by the intervention. The bibasic vector is embedded into a predefined three-dimensional event space. The bibasic vector in the three-dimensional event space is detected by a geometric constraint model of the vehicle's physical state. Inconsistent event units are identified and corrected, generating a set of corrected vectors to eliminate the inconsistencies. The corrected vector set is then projected onto a predefined used car valuation model data field space to generate structured data units aligned with the used car valuation model data field space. These structured data units are used to construct a standardized used car dataset.
[0004] On the other hand, embodiments of the present invention provide a dataset construction apparatus, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the program to implement the steps in the above-described method.
[0005] Thirdly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the methods described above.
[0006] The present invention provides a data governance-based method for constructing used car datasets. After cleaning and standardizing the original vehicle description text to obtain a set of vehicle description texts to be processed, it further extracts event units containing operational behaviors and state changes. This method can accurately capture the inherent correlation between operational intervention and state evolution in different stages of vehicle use, avoiding the semantic fragmentation problem caused by isolated extraction of text fragments or single attributes in traditional data processing. By performing bibasic vector decomposition on each event unit, the event unit is transformed into a bibasic vector composed of operation vectors and state vectors. This realizes the transformation of qualitative events described in natural language into quantifiable mathematical representations, enabling precise characterization of the intervention intensity of operational behaviors and the offset magnitude of state changes, breaking through the limitations of traditional text processing. This vectorization ignores the limitations of causal relationships; it embeds bibasic vectors into a three-dimensional event space and detects and corrects them through a geometric constraint model of the vehicle's physical state. This enables the identification and elimination of contradictions in event units based on the vehicle's physical laws, ensuring that the data conforms to the objective physical logic of vehicle use and avoiding subsequent analysis biases caused by data contradictions. Finally, through projection transformation, the corrected vector set is mapped to the data field space of the used car valuation model, generating aligned structured data units. This allows the constructed dataset to directly adapt to the field requirements of downstream valuation models, avoiding the problem of low data utilization caused by the disconnect between data governance and business applications. This improves the semantic integrity, quantitative accuracy, physical rationality, and business adaptability of the used car dataset as a whole. Attached Figure Description
[0007] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present invention and, together with the specification, serve to explain the technical solutions of the present invention.
[0008] Figure 1 This is a schematic diagram illustrating the implementation process of a data governance-based used car dataset construction method provided in an embodiment of the present invention.
[0009] Figure 2 This is a schematic diagram of a hardware entity of a dataset construction device provided in an embodiment of the present invention. Detailed Implementation
[0010] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on the present invention. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0011] An embodiment of the present invention provides a method for constructing a used car data set based on data governance, which can be executed by a data set construction device. The data set construction device may refer to a device with data processing capabilities such as a server, a laptop computer, a tablet computer, a desktop computer, etc.
[0012] Figure 1 As shown in the schematic diagram of the implementation process of a method for constructing a used car data set based on data governance provided by an embodiment of the present invention, Figure 1 as shown, the method includes: Step S100: Obtain the original vehicle description text in the used car transaction scenario, perform text cleaning and standardization processing on the original vehicle description text to obtain a set of vehicle description texts to be processed, and the set of vehicle description texts to be processed includes status description information and operation record information of the vehicle at different usage stages.
[0013] The original vehicle description text comes from various channels in the used car transaction scenario, such as vehicle introductions on online trading platforms, records in vehicle maintenance manuals, etc. These texts may contain a large amount of noise information, such as non-standard expressions, typos, special symbols, etc. At the same time, the text formats and word usages may also vary.
[0014] To obtain the original vehicle description text, web crawler tools can be used to capture relevant text content on major used car trading websites. For offline paper documents such as maintenance manuals, they can be converted into electronic texts through optical character recognition (OCR) technology. It can be understood that the captured information is publicly available information in public channels and complies with laws and regulations.
[0015] During text cleaning, first use regular expressions to remove special symbols and extra spaces in the text. Regular expressions are a powerful text matching tool. By defining specific patterns, unnecessary characters can be accurately identified and removed. Then, a spelling check model based on machine learning is used to correct typos. This model is trained with a large amount of correct text data to learn common spelling error patterns and correct errors based on context information. For example, if "maintenance" is written as "maintenance sample" in the text, the model can correct it to the correct expression.
[0016] Standardization processing is to unify the expressions in the text into standard vocabulary and formats. Establish a standard vocabulary list covering common professional terms and standard expressions in the used car field, and replace non-standard vocabulary in the text with standard vocabulary. For example, "performed maintenance" is unified as "carried out maintenance", and "the tires were changed" is expressed as "replaced the tires". After cleaning and standardization processing, the vehicle description text to be processed "The vehicle had maintenance last year, the tires were replaced, there are scratches on the exterior, and the price is negotiable" is obtained.
[0017] The set of vehicle description texts to be processed consists of multiple processed vehicle description texts. The status description information includes the vehicle's appearance, performance, and other aspects of its condition, while the operation record information covers operations such as repair, maintenance, and parts replacement. For example, the set may also contain the text "The vehicle was in a collision the year before last, was repaired and is now in normal use, and the spark plugs were replaced this year."
[0018] Step S200: Extract event units from the set of vehicle description texts to be processed. Identify and extract event units containing operational behaviors and state changes in the set of vehicle description texts to be processed. Event units are used to characterize the state change process of the vehicle at the corresponding time node.
[0019] Event unit extraction aims to identify the parts describing vehicle operation behaviors and their resulting state changes from the set of vehicle description texts to be processed, and organize them into units with clear logical relationships. Operation behaviors are various actions performed on the vehicle, such as maintenance, repair, and replacement of parts; state changes are the different states that the vehicle exhibits after operation behaviors, such as performance improvements and appearance changes.
[0020] For example, in the vehicle description text "maintenance was performed last year, and vehicle performance was improved", "maintenance was performed" is an operation, and "vehicle performance was improved" is a state change. The two constitute an event unit, representing the state change process in which the vehicle's performance was improved after the operation of maintenance was performed last year.
[0021] In one implementation, step S200 may specifically include the following steps S210 to S260: Step S210: Perform time stamp parsing on the set of vehicle description texts to be processed, extract the descriptive information representing time nodes in the text, divide the set of vehicle description texts to be processed into text fragment units with sequential order based on the time node descriptive information, perform semantic coherence analysis on adjacent text fragment units, calculate the semantic overlap between the end of the preceding fragment and the beginning of the following fragment, and complete the content of text fragments whose semantic overlap does not meet the set standard.
[0022] Time stamp parsing involves extracting time-related words or phrases from the vehicle description text, such as "last year," "early this year," and "the year before last." These time stamps are crucial for determining the chronological order of events. By extracting time stamps, the text can be arranged and divided chronologically.
[0023] For example, for the set of vehicle description text to be processed, “The vehicle was maintained last year, and its performance was improved; the tires were replaced at the beginning of this year, and its driving stability was enhanced; the vehicle was involved in a collision the year before last,” the time markers “last year,” “the beginning of this year,” and “the year before last” are extracted, and the text is divided into three text segment units based on these: “The vehicle was maintained last year, and its performance was improved”, “The tires were replaced at the beginning of this year, and its driving stability was enhanced”, and “The vehicle was involved in a collision the year before last.”
[0024] Semantic coherence analysis ensures semantic coherence between adjacent text segments, preventing logical jumps or irrelevant information. Semantic overlap is measured by the proportion of identical words or phrases in adjacent text segments. A bag-of-words model can be used to convert text segments into vector representations, and then cosine similarity can be calculated between the vectors; the closer the cosine similarity is to 1, the more semantically similar the segments.
[0025] If the set semantic overlap standard is a certain level, when the semantic overlap between "maintenance was performed last year, improving vehicle performance" and "tires were replaced at the beginning of this year, enhancing driving stability" does not meet the standard, the content needs to be supplemented by combining contextual information or domain knowledge. For example, it can be supplemented as "the vehicle was used normally after maintenance last year, and the tires were replaced at the beginning of this year to improve driving stability."
[0026] Step S220: Perform semantic role recognition on the text fragment units, analyze the action initiator, action target and state change description in each text fragment unit, and determine the specific vehicle component or operation execution subject pointed to by the pronoun in combination with the contextual reference relationship, so as to obtain the annotated text content containing semantic role identifiers and reference relationships.
[0027] Semantic role recognition determines the role of each word in a text within an event description. The action initiator is the entity performing the action, the action target is the object to which the action is directed, and the state change description describes the changes in vehicle state caused by the action.
[0028] For example, in the text fragment "Last year, maintenance personnel performed maintenance on the vehicle, and the vehicle's performance was improved", "maintenance personnel" is the initiator of the action, "vehicle" is the object of the action, and "vehicle performance improved" is the description of the state change.
[0029] Determining the specific vehicle component or operational entity referred to by a pronoun by considering its contextual referential relationships can eliminate referential ambiguity in the text. For example, in the text "The car is undergoing repairs, and it is currently in good condition," "it" refers to "the car."
[0030] In one implementation, step S220 may specifically include the following steps S221 to S226: Step S221: Segment the text fragment units into words, dividing the continuous text sequence into word components with independent semantic functions. The word components are divided into different categories according to their part-of-speech features, and each category of words undertakes different semantic functions in the event description.
[0031] Word segmentation is the process of dividing text fragments according to word boundaries to obtain independent words. In Chinese text processing, tools such as Jieba segmentation can be used. Jieba segmentation, based on statistical models and rules, can accurately segment Chinese text into words.
[0032] For example, the text fragment "Last year, the vehicle's performance was improved" can be segmented using Jieba to obtain the word components "last year", "performed", "maintenance", "vehicle", "performance", and "improvement".
[0033] Word components are classified according to their parts of speech. Common parts of speech include nouns, verbs, adjectives, and adverbs. Words of different parts of speech play different semantic functions in the description of events. Nouns usually indicate the object or related entity of the action, verbs indicate the operation, and adjectives and adverbs are used to describe the state or modify the action.
[0034] Step S222: Call the context semantic coding model to perform semantic association processing on word components, analyze the positional relationship of word components in the text segment and the semantic influence of adjacent words, and generate a word vector sequence containing context-dependent information. The word vector sequence reflects the semantic connotation of word components in a specific context.
[0035] Contextual semantic encoding models can take into account the semantic information of words in their context, such as the BERT model. BERT is based on the Transformer architecture and learns the rich semantic information of words in their context through a bidirectional attention mechanism.
[0036] Word components are input into the BERT model. The model encodes each word based on its position in the text segment and information from adjacent words, generating a corresponding word vector. For example, for the word "maintenance," in the context of "maintenance was performed last year, and the vehicle's performance was improved," its vector representation would contain semantic information related to "vehicle" and "performance improvement."
[0037] Step S223: Input the word vector sequence into the semantic role recognition process, perform role classification processing on each word component, and output the semantic role classification result corresponding to each word component. The classification result reflects the functional positioning of the word component in the event description.
[0038] The semantic role recognition process is based on machine learning algorithms, such as support vector machines or neural network models. These models are trained on a large amount of labeled data to learn word feature patterns corresponding to different semantic roles.
[0039] Input the sequence of word vectors into the trained semantic role recognition model, and the model classifies each word component. For example, "perform" is classified as the "action initiator" role, "vehicle" is classified as the "object of action" role, and "performance improvement" is classified as the "state change" role.
[0040] Step S224: Determine the semantic role label for each word component based on the semantic role classification result. The semantic role label reflects the functional position of the word component in the event description, covering functions such as action initiator, object of action, state change, and background description.
[0041] According to the semantic role classification result, determine the specific semantic role label for each word component. For example, the semantic role label of "last year" is "background description", "perform" is "action initiator", "vehicle" is "object of action", and "performance improvement" is "state change".
[0042] Step S225: Store the semantic role label in association with the corresponding word component, establish the correspondence between the word component and the semantic role label, and obtain an associated data group containing the word component,词性类别 (not sure what this means exactly, might be "word class" or something similar), and semantic role label.
[0043] Store each word component in association with its corresponding semantic role label, which can be implemented using a database or a data structure. For example, using a dictionary data structure, store the word component as the key and the semantic role label as the value, and at the same time store the词性类别信息 (not sure what this means exactly, might be "word class information" or something similar).
[0044] For the above example, the associated data group can be represented as: {"last year":["time noun","background description"],"perform":["verb","action initiator"],"maintenance":["noun","related to action initiation"],"vehicle":["noun","object of action"],"performance":["noun","related to object"],"improvement":["verb","state change"]}.
[0045] Step S226: Conduct cross-checking on the associated data group, compare the common semantic role configurations in the field, and adjust the deviated parts in the recognition result to make the associated data group conform to the semantic norms in the used car field, and obtain the annotated text content.
[0046] Cross-checking ensures that the semantic role labels in the associated data group are accurate and conform to the semantic norms in the used car field. Collect a large amount of annotated data in the used car field, count the common semantic role configurations, and form a reference data set.
[0047] The associated data set is compared with the reference dataset. If some semantic role identifiers in the identification results do not match the common configurations in the reference dataset, adjustments are made. For example, in the used car industry, "maintenance" is usually considered a specific operation on the vehicle, and its semantic role identifier should be more clearly classified as "action initiation related." If the previous identification results were biased, they need to be corrected.
[0048] After adjustment, the annotated text content containing accurate semantic role identification and referential relationships is obtained, providing a reliable foundation for the subsequent construction of event units.
[0049] Step S230: Based on the association between the action initiator and the action target in the labeled text content, construct the basic compositional relationship of vehicle operation behavior, analyze the temporal continuity characteristics of the action, and distinguish between one-time operation and continuous operation.
[0050] The basic structural relationships of vehicle operation behavior refer to the logical relationships between the initiator of the action, the object of the action, and the operation itself. By analyzing the association between the initiator and the object of the action in the annotated text, the specific components of the operation behavior can be clarified.
[0051] For example, in the text “Last year, maintenance personnel performed maintenance on the vehicle”, “maintenance personnel” is the initiator of the action, “vehicle” is the object of the action, and “maintenance” is the operation, which constitutes the basic relationship of “maintenance personnel-maintenance-vehicle”.
[0052] In terms of the time duration characteristics of actions, one-time operations are those completed in a short period of time, such as "changing a tire" or "changing brake fluid"; continuous operations are those performed over a period of time, such as "a vehicle parked for a long time" or "a vehicle continuously driving".
[0053] Distinguish between one-off and continuous operations by examining the time descriptions and the nature of the actions in the text. For example, "The tires were replaced last year" is a one-off operation, as the "replacement" was completed in one go; "The vehicle has been parked for a period of time" is a continuous operation, as the "parking" was ongoing.
[0054] Step S240: Based on the semantic association analysis results of the relationship between the state change description and the basic composition of vehicle operation behavior, determine the state change description content corresponding to each operation behavior, and combine the causal relationship of vehicle attributes to identify the main state changes directly caused by the operation behavior and the related state changes indirectly caused by it, so as to obtain a preliminary set of event units containing the correspondence between operation behavior and state change.
[0055] Semantic association analysis identifies the correspondence between the state change description and the basic components of vehicle operation behavior by comparing semantic information in these two parts. For example, in the annotated text "Vehicle performance improved after maintenance last year", semantic association analysis determines that the state change description corresponding to the "maintenance" operation is "vehicle performance improved".
[0056] Causal relationships in vehicle attributes refer to the causal connections between various vehicle attributes. For example, tire wear directly affects driving stability, which in turn indirectly affects fuel consumption. By analyzing these causal relationships, we can more accurately identify the primary state changes directly caused by operational behaviors and the related state changes indirectly resulting from them.
[0057] For the operation of "changing tires", the main direct change in state may be "enhanced driving stability", and the related indirect change may be "reduced fuel consumption".
[0058] Each operation and its corresponding state change description are associated to form an event unit, resulting in a preliminary set of event units. For example, the set might include event units such as "maintenance performed last year - vehicle performance improved" or "tire replaced last year - improved driving stability and reduced fuel consumption."
[0059] Step S250: Redundancy check is performed on the initial set of event units. Based on the combination characteristics of operation behavior type, target, and state change amplitude, the redundancy between event units is calculated. Event units with redundancy reaching the set standard are removed to obtain the deduplicated event unit group.
[0060] Redundancy checks remove duplicate event units from the initial set of event units to avoid data redundancy and interference. The combination of operational behavior type, target object, and state change magnitude is used to measure the similarity of event units.
[0061] Define a repeatability calculation function to calculate the repeatability between two event units based on a combination of characteristics including the type of operation, the target object, and the magnitude of state changes. If a certain repeatability standard is set, one of the event units is removed when its repeatability calculation result reaches or exceeds that standard.
[0062] By comparing each pair of event units in the initial set of event units and calculating the degree of repetition, a group of event units after deduplication is obtained.
[0063] Step S260: Dependency identification is performed on the deduplicated event unit group, the causal and conditional relationships between event units are analyzed, and the preceding event identification information and the following event identification information are added to each event unit to obtain a structured event unit containing time sequence, semantic role and dependency relationship.
[0064] Causal association refers to the occurrence of one event unit leading to the occurrence of another event unit, such as "vehicle collision - vehicle exterior damage," where "vehicle collision" is the cause and "vehicle exterior damage" is the result. Conditional association refers to the occurrence of one event unit requiring the fulfillment of a condition for another event unit, such as "vehicle reaches a certain mileage - maintenance is required," where "vehicle reaches a certain mileage" is the condition for "maintenance is required." Analyzing the logical relationships between event units in the deduplicated event unit group helps determine causal and conditional associations. For each event unit, pre-event and post-event identifiers are added. Pre-event identifiers indicate events that must occur before the current event unit, while post-event identifiers indicate events that may occur after the current event unit.
[0065] For example, for the event unit "maintenance performed last year - vehicle performance improved", if its preceding event is "vehicle traveled a certain mileage", then add the preceding event label "vehicle traveled a certain mileage"; if the following event is "vehicle can drive more smoothly", then add the following event label "vehicle can drive more smoothly".
[0066] After processing, structured event units containing time sequence, semantic roles, and dependencies are obtained, providing a clear logical structure for subsequent vector decomposition and dataset construction.
[0067] Step S300: Perform bibasic vector decomposition on each event unit, decomposing the event unit into bibasic vectors composed of operation vectors and state vectors. The operation vectors are used to characterize the force and direction of the active intervention applied to the vehicle, and the state vectors are used to characterize the offset magnitude of the vehicle attributes under the intervention.
[0068] Bibasic vector decomposition transforms event units into operation vectors and state vectors, facilitating the quantification and analysis of event units. Operation vectors reflect the force and direction of vehicle operations; for example, in the operation of "changing a tire," the operation vector can represent information such as the speed and force of tire changing. State vectors reflect the changes in the vehicle's state after the operation, such as changes in vehicle driving performance and appearance.
[0069] For example, the event unit "Tire replacement last year - enhanced driving stability" can be decomposed into an operation vector and a state vector. The operation vector contains information such as the force and time of the tire replacement operation, while the state vector contains information such as the magnitude of the improvement in driving stability and changes in driving smoothness.
[0070] In one implementation, step S300 may specifically include the following steps S310 to S360: Step S310: Perform a structured transformation on the event unit, converting the event unit in natural language form into a structured event description consisting of an operational behavior description, an explanation of the target object, operational implementation tools, environmental influencing factors, and state changes. Each component has a clear logical relationship in the event description.
[0071] Structured transformation converts event units described in natural language into fixed-structured expressions, facilitating subsequent processing and analysis. The operational behavior description clearly specifies the concrete operation performed on the vehicle, such as "changing a tire"; the target object description indicates the object to which the operation is performed, i.e., the "vehicle"; the operation tools refer to the tools used during the operation, such as "tire wrench"; environmental influencing factors consider the environmental conditions at the time of the operation, such as "indoor repair shop"; and the state change description depicts the changes in the vehicle's state caused by the operational behavior, such as "enhanced driving stability".
[0072] For example, for the event unit "Last year in the indoor repair shop, maintenance personnel used a tire wrench to change the tire of a vehicle, which enhanced driving stability", the transformed structured event description is as follows: Operation behavior description - "Tire change"; Object description - "Vehicle"; Operation tool - "Tire wrench"; Environmental influencing factors - "Indoor repair shop"; Change in state - "Enhanced driving stability".
[0073] Step S320: Extract emerging terms from the description of vehicle operation behavior and state changes, add the extracted emerging terms to the vehicle operation behavior vocabulary set and the vehicle state attribute vocabulary set, and update the term entries and semantic descriptions of the vocabulary sets.
[0074] Emerging terminology extraction identifies new terms not included in the existing vocabulary that appear in descriptions of vehicle operation behavior and state changes. With the development of the used car industry and technological advancements, new terms describing operation behavior and state may emerge.
[0075] For example, when describing vehicle operating behavior, the emerging term "intelligent tire replacement" may appear; when describing changes in state, the term "vehicle intelligent system performance optimization" may appear.
[0076] The extracted emerging terms are added to the vehicle operation behavior vocabulary set and the vehicle state attribute vocabulary set, and the term entries and semantic descriptions in the vocabulary sets are updated. Emerging term extraction can be performed using rule-based methods or machine learning methods. Rule-based methods identify new terms by defining specific patterns, while machine learning methods use named entity recognition models, such as BiLSTM-CRF-based models, which learn from large amounts of text data to accurately identify emerging terms in the text.
[0077] Step S330: Based on the updated vehicle operation behavior vocabulary set, perform joint vectorization processing on the operation behavior descriptions and operation implementation tools in the structured event representation. Combining the type characteristics of operation behavior and the functional characteristics of operation implementation tools, generate a comprehensive operation vector that integrates operation type, implementation intensity and tool influence. The dimension of the comprehensive operation vector matches the terminology scale of the updated vehicle operation behavior vocabulary set.
[0078] Joint vectorization processing converts the descriptions of operational behaviors and the information on the tools used to perform operations in a structured event representation into vector representations. The type characteristics of the operational behavior include the nature of the operation, such as maintenance, repair, and replacement; the functional characteristics of the operational tools include the tool's purpose and efficiency.
[0079] First, word embedding techniques are used to convert the textual information of the operational behavior description and the operational implementation tool into vectors. For example, Word2Vec or GloVe can be used to convert words into vector representations. Then, the vectors are fused by combining the type characteristics of the operational behavior and the functional characteristics of the operational implementation tool. The operational behavior type characteristic vector and the operational implementation tool functional characteristic vector are fused using a weighted summation method.
[0080] Next, based on the intensity modifiers in the description of the operational behavior and the analysis results of the tool function parameters, the intensity level of the operation is determined. Intensity modifiers, such as "quick change" and "slow adjustment," combined with the analysis results of tool function parameters, such as tool power and accuracy, determine the intensity level of the operation. The intensity level is then converted into an intensity coefficient that matches the basic vector dimension.
[0081] Finally, based on the mapping relationship between tool function parameters and operational effects, different weights are assigned to each dimension of the comprehensive vector. The mapping relationship between tool function parameters and operational effects can be established through experimental or empirical data. The comprehensive vector is multiplied element-wise with the weight vector to generate a comprehensive operational vector that integrates operation type, implementation intensity, and tool impact, with its dimensions matching the terminology size of the updated vehicle operation behavior vocabulary set.
[0082] In one implementation, step S330 may specifically include the following steps S331 to S336: Step S331: Perform principle analysis on the operational behavior descriptions in the structured event representation. Based on the differences in the effects of the operation on the vehicle's physical structure or electronic system, classify the operational behavior into different types of effects and record the proportion of each category in the operational behavior description.
[0083] The principle analysis of operational behavior provides a deeper understanding of how operational behavior affects the vehicle in specific ways. Based on the differences in how operations affect the vehicle's physical structure or electronic systems, operational behavior can be categorized into different types of actions, such as mechanical operations, electronic operations, and maintenance operations.
[0084] For example, "changing a tire" is a mechanical operation, while "upgrading a vehicle software system" is an electronic operation. Statistics are compiled showing the percentage of each category in the descriptions of operational behaviors to understand the distribution of different types of operational behaviors.
[0085] Step S332: Deconstruct the operation implementation tools in the structured event description, extract the descriptive information of the tool's functional components, operation methods and accuracy levels, and establish a mapping relationship between tool function parameters and operation effects.
[0086] The functional breakdown of the operating tool involves a detailed analysis of various aspects of the tool. The working part is the part of the tool that directly participates in the operation, such as the head of a "tire wrench"; the operation method is the way to use the tool, such as "rotation tightening"; the accuracy level description information reflects the operating precision of the tool, such as "higher precision".
[0087] By analyzing the functional parameters of operational tools, a mapping relationship between tool functional parameters and operational effects can be established. For example, the higher the precision of a tire wrench, the better the operational effect when changing tires and the higher the tire installation stability. This mapping relationship can be established through experimental or empirical data, linking the tool's functional parameters with quantitative indicators of operational effects.
[0088] Step S333: Construct a dynamic operation dictionary based on the updated vehicle operation behavior vocabulary set, map operation behaviors of different function types to corresponding entries in the dictionary, and generate a basic vector representing the operation type.
[0089] The dynamic operation dictionary is constructed based on the updated vocabulary set of vehicle operation behaviors, mapping different types of operation behaviors to entries in the dictionary. For example, "changing tires" is mapped to the dictionary entry "mechanical operation - tire changing".
[0090] A basic vector is generated for each entry. Using one-hot encoding, each operation type corresponds to a vector, where only one element has a specific value, and the rest have another specific value. For example, the basic vector corresponding to "Mechanical Operation - Tire Change" has a specific form.
[0091] Step S334: Based on the intensity modifiers in the description of the operation behavior and the analysis results of the tool function parameters, determine the intensity level of the operation and convert the intensity level into an intensity coefficient that matches the basic vector dimension.
[0092] Intensity modifiers refer to words used to describe the force of an operation in the description of the action, such as "quick change" or "slow adjustment." The intensity level of the operation is determined by combining the analysis results of the tool's functional parameters.
[0093] For example, "quick tire change" indicates a relatively large operating force, and its force level is determined to be "high" based on the strength modifier and tool function parameters. The force level is then converted into a strength coefficient that matches the underlying vector dimension.
[0094] Step S335: Element-wise merge the base vector and intensity coefficient to obtain an intermediate vector containing operation type and implementation intensity. The dimension of the intermediate vector is consistent with the number of entries in the dynamic operation dictionary.
[0095] The intermediate vector is obtained by element-wise multiplying the base vector with the intensity coefficients. The intermediate vector contains information on the operation type and intensity, and its dimension matches the number of entries in the dynamic operation dictionary, ensuring accurate representation of combinations of different operation types and intensities.
[0096] Step S336: Based on the mapping relationship between tool function parameters and operation effects, assign differentiated weights to each dimension of the intermediate vector, and generate a comprehensive operation vector that integrates operation type, implementation intensity and tool influence through weighted calculation.
[0097] Based on the mapping relationship between tool function parameters and operational effects, different weights are assigned to each dimension of the intermediate vector. For example, the accuracy of the "tire wrench" has a significant impact on the operational effect, so the dimensions related to "tire changing" in the intermediate vector are assigned higher weights.
[0098] The intermediate vector and the weight vector are multiplied element-wise and then summed to generate a comprehensive operation vector that integrates operation type, implementation strength, and tool influence.
[0099] Step S340: Based on the updated vehicle state attribute vocabulary set, perform joint vectorization processing on the state changes and environmental influencing factors in the structured event description, distinguish between recoverable and unrecoverable state changes, add a continuous impact weight to unrecoverable state changes, and generate a comprehensive state vector that includes the magnitude of state attribute changes and the degree of environmental impact. The dimension of the comprehensive state vector matches the terminology scale of the updated vehicle state attribute vocabulary set.
[0100] Joint vectorization processing converts the state changes and environmental influencing factors in structured event representations into vector representations. Recoverable state changes refer to changes in which a vehicle can return to its original state after a certain operation or time, such as "minor scratches on the vehicle's exterior"; irrecoverable state changes refer to changes that cannot be fully recovered, such as "severe damage to the vehicle's engine".
[0101] We differentiate between recoverable and irreversible state changes, adding a weighted approach to irreversible state changes to reflect their long-term impact on vehicle status. We use word embedding technology to convert terms related to state changes and environmental factors into vectors, then perform a weighted summation. Considering the impact of environmental factors on vehicle status, such as "vehicle performance degradation under high temperature conditions," we calculate the contribution ratio of environmental factors to state changes.
[0102] The integrated state vector dimension matches the term size of the updated vehicle state attribute vocabulary set, accurately representing various attribute changes and environmental influences of the vehicle state.
[0103] As one implementation method, step S340 involves jointly vectorizing the state changes and environmental influencing factors in the structured event description based on the updated vehicle state attribute vocabulary set, distinguishing between recoverable and irrecoverable state changes, adding a continuous impact weight to irrecoverable state changes, and generating a comprehensive state vector that includes the magnitude of state attribute changes and the degree of environmental impact. Specifically, this may include the following steps S341~S346: Step S341: Perform attribute tracing on the state changes in the structured event description, determine the core vehicle attributes involved in the state changes, and establish the association and correspondence between the state change description and the core attributes.
[0104] Attribute tracing identifies the core vehicle attributes involved in the state change description. These core attributes significantly impact vehicle performance and value, such as engine performance and tire wear. For example, for the state change description "enhanced driving stability," analysis determines that the relevant core vehicle attributes are tire performance and suspension system performance. Establish a correlation between the state change description and these core attributes, such as associating "enhanced driving stability" with "improved tire performance" and "optimized suspension system performance."
[0105] When tracing attribute origins, one can first perform semantic analysis on the descriptions of state changes to identify key descriptive terms, and then combine this with vehicle domain knowledge to determine the core attributes corresponding to these terms. For example, "driving stability" is related to attributes such as tires and suspension systems. By further analyzing the reasons or manifestations of enhanced driving stability in the text, the specific core attributes involved can be determined.
[0106] Step S342: Perform an impact analysis on the environmental factors in the structured event description, and calculate the contribution ratio of each factor to the state change based on the description information of the duration and scope of the environmental factors.
[0107] The analysis assesses the impact of environmental factors on changes in vehicle condition. Duration refers to the length of time the environmental factor exists, and the scope of impact describes the vehicle components or systems affected by the environmental factor.
[0108] For example, regarding the phrase "vehicle engine performance declines under high-temperature conditions," the longer the high-temperature environment lasts, the greater the impact on engine performance. Based on the description of the duration and scope of the environmental factors, the contribution ratio of each factor to the state change is calculated. Specifically, when calculating the contribution ratio, historical data can be collected first to analyze the changes in vehicle state under different environmental conditions. For each environmental factor, the degree of change in vehicle state under different durations and scopes is statistically analyzed. Then, based on the current duration and scope of the environmental factor, and referring to historical data, the contribution ratio of that environmental factor to the state change is determined. For example, if historical data shows that when a high-temperature environment lasts for a certain duration and affects the engine, the degree of engine performance decline is proportional to the duration, then the contribution ratio to the engine performance decline can be calculated proportionally based on the current duration of the high-temperature environment.
[0109] Step S343: Construct an attribute feature space based on the updated vehicle state attribute vocabulary set, map the associated core attributes to the corresponding dimensions in the feature space, and generate a basic vector representing the state attributes.
[0110] The attribute feature space is a multi-dimensional space, with each dimension corresponding to a vehicle state attribute. The attribute feature space is constructed based on the updated vehicle state attribute vocabulary, mapping the associated core attributes to their corresponding dimensions within the feature space.
[0111] For example, "engine performance" can be mapped to a specific dimension of the feature space. A basic vector representing the state attributes is generated, with each element corresponding to a core attribute. If the feature space has multiple dimensions, the dimension corresponding to the associated core attribute is set to a specific value, and the remaining dimensions are set to another specific value, resulting in the basic vector.
[0112] Step S344: Based on the change amount description in the state change description and the contribution ratio of environmental factors, calculate the change magnitude value of each core attribute, and convert the change magnitude value into an amplitude coefficient that matches the basic vector dimension.
[0113] The change quantity describes the specific degree of state change, such as "driving stability has improved to a certain extent." The magnitude of change for each core attribute is calculated by considering the contribution ratio of environmental factors.
[0114] For example, regarding "enhanced driving stability," considering the influence of environmental factors, the changes in "tire performance" and "suspension system performance" are calculated. These changes are then converted into amplitude coefficients that match the dimensions of the base vector, which can be done using certain conversion rules.
[0115] Step S345: Determine the reversibility of the state change, distinguish between recoverable and unrecoverable state changes, add time decay weights to the amplitude coefficients corresponding to unrecoverable state changes, and generate weighted amplitude coefficients.
[0116] To determine the reversibility of a state change, the amplitude coefficient remains constant for recoverable state changes; for irreversible state changes, a time decay weight is added. The time decay weight takes into account that the impact of irreversible state changes gradually weakens over time, but still has a certain lasting effect. For example, for an irreversible state change such as "severe engine damage," a time decay weight is added. Multiplying the amplitude coefficient by the time decay weight yields the weighted amplitude coefficient.
[0117] Step S346: Element-wise merge the base vector and the weighted amplitude coefficients to obtain an intermediate vector containing state attributes and change amplitudes. The dimension of the intermediate vector is consistent with the dimension of the attribute feature space. Adjust the values of each dimension of the intermediate vector according to the contribution ratio of environmental factors to generate a comprehensive state vector.
[0118] The base vector is multiplied element-wise with the weighted amplitude coefficients to obtain an intermediate vector that contains state attributes and their magnitudes of change. The intermediate vector has the same dimension as the attribute feature space, accurately representing the changes in vehicle state attributes.
[0119] The values of each dimension of the intermediate vector are adjusted according to the contribution ratio of environmental factors. If an environmental factor contributes a large proportion to a certain core attribute, the value of that dimension is increased accordingly. After adjustment, a comprehensive state vector is generated, which includes information on the magnitude of changes in vehicle state attributes and the degree of environmental impact.
[0120] Step S350: Perform causal correlation analysis on the integrated operation vector and the integrated state vector, statistically analyze the co-occurrence of operation behavior and state change in historical event data, determine the degree of correlation between the two, filter vector pairs with a correlation degree greater than the first threshold, and exclude abnormal vector pairs with a correlation degree less than the second threshold.
[0121] Causal correlation analysis determines the strength of the causal relationship between the integrated operation vector and the integrated state vector. By statistically analyzing the co-occurrence of operational behaviors and state changes in historical event data, the degree of correlation between them is understood.
[0122] For example, count the number of times the "tire replacement" operation and the "improved driving stability" status change occur in historical data. If the number of co-occurrences is high, it indicates that the two are closely related.
[0123] Set a first threshold and a second threshold to filter vector pairs whose correlation is greater than the first threshold. These vector pairs indicate a strong causal relationship between the operation and the state change. Exclude abnormal vector pairs whose correlation is less than the second threshold, as these pairs may have weak correlations due to data errors or other anomalies.
[0124] The degree of correlation between the integrated operation vector and the integrated state vector can be calculated using methods such as the Pearson correlation coefficient, and the results can be used for filtering and exclusion.
[0125] Step S360: Perform dimensionality optimization on the integrated operation vector and integrated state vector, retain key information, maintain the causal relationship between operation and state, reduce the vector dimension, reduce redundant information, and obtain a bibasic vector composed of the operation vector and state vector.
[0126] Dimensionality optimization reduces the dimensions of the integrated operation vector and integrated state vector while retaining key information and operation-state causal relationships. Principal component analysis can be used for dimensionality optimization.
[0127] Principal component analysis (PCA) reduces data dimensionality by identifying principal components. The integrated operation vector and integrated state vector are input into the PCA model. The model calculates the covariance matrix of the data, solves for the eigenvalues and eigenvectors of the covariance matrix, selects principal components with larger eigenvalues, and projects the data onto these principal components to obtain the dimensionality-reduced vectors.
[0128] During dimensionality reduction, it is crucial to preserve the causal relationships between operations and states. This can be achieved by adding constraints to the principal component analysis model or by using an improved principal component analysis method. After dimensionality optimization, a bibasic vector consisting of operation and state vectors is obtained. These vectors are more concise and facilitate subsequent processing and analysis.
[0129] Step S400: Embed the bibasic vectors into the preset three-dimensional event space, detect the bibasic vectors in the three-dimensional event space through the geometric constraint model of the vehicle's physical state, identify and correct contradictory event units, and generate a set of corrected vectors to eliminate contradictions.
[0130] Bistudent vectors are embedded into a predefined three-dimensional event space to analyze the relationship between operation vectors and state vectors in three-dimensional space. The three dimensions of the three-dimensional event space are time dimension, operation intensity dimension, and state change dimension.
[0131] The geometric constraint model of the vehicle's physical state is established based on the vehicle's physical characteristics and operational logic, defining the geometric constraints that the bibasic vectors should satisfy in the three-dimensional event space. For example, there should be a certain causal relationship between operational intensity and state changes, and events should occur sequentially in the time dimension.
[0132] This geometric constraint model is used to detect bibasic vectors in the three-dimensional event space and identify contradictory event units. Contradictions may manifest as a mismatch between operation intensity and state change, or an illogical order of event occurrence.
[0133] For event units with contradictions, the corresponding correction strategy is invoked according to the contradiction type to adjust the parameters of the basis vectors so that the corrected basis vectors satisfy the conditions of the geometric constraint model, thereby generating a set of corrected vectors that eliminates the contradictions.
[0134] In one implementation, step S400 may specifically include the following steps S410 to S460: Step S410: Determine the dimensional composition of the three-dimensional event space. The time dimension includes the absolute time marker of the event occurrence and the time interval between adjacent events. The operation intensity dimension includes the force of the operation and the frequency of the operation within a unit time period. The state change dimension includes the immediate state performance after the operation and the cumulative state performance. Each dimension converts the original values into standardized coordinate values within a unified range through linear transformation.
[0135] The three-dimensional event space is structured with clearly defined dimensions, each with its own meaning. The time dimension reflects the timing of events; absolute time markers determine the specific point in time when an event occurs, and the time interval between adjacent events reflects their temporal sequence. The operation intensity dimension measures the strength and frequency of the operation; the force of the operation indicates its magnitude, and the frequency of operations within a unit of time indicates the number of times the operation occurs within a certain period. The state change dimension describes the changes in vehicle state caused by the operation; the immediate state after the operation is the state change immediately present after the operation, while the cumulative state is the accumulated change in vehicle state after multiple operations.
[0136] To facilitate comparison and analysis in a three-dimensional event space, the original numerical values of each dimension are converted into standardized coordinate values within a unified interval using a linear transformation method. The linear transformation maps the original numerical values to a fixed interval, such as a specific range, through a linear transformation.
[0137] When performing a linear transformation, first determine the range of values for the original values in each dimension, and then calculate the transformation coefficients based on this range and the target standardized interval. For example, if the range of the original values for a dimension is from the minimum to the maximum value, and the target standardized interval is from a specific lower limit to a specific upper limit, the original values can be converted into standardized coordinate values by calculating the difference ratio.
[0138] In one implementation, step S410 may specifically include the following steps S411 to S416: Step S411: Divide the time dimension into two coordinates. The absolute time stamp is converted into a time series value by parsing the time description information of the event unit. The time interval between adjacent events is generated by calculating the difference between the time stamps of the current event and the previous event, thus obtaining two sub-coordinate components of the time dimension.
[0139] The dual-coordinate partitioning of the time dimension subdivides the time dimension into two sub-coordinate components: absolute time stamps and time intervals between adjacent events. For absolute time stamps, the time description information in the event unit, such as "last year" or "early this year," is parsed and converted into specific time series values. These values can be represented using timestamps, which are numbers representing specific points in time.
[0140] For example, convert "last year" to a specific year value, and calculate the timestamp corresponding to "last year" based on the current year. The time interval between adjacent events is obtained by calculating the difference between the timestamps of the current event and the preceding event. If the timestamp of the current event is a specific time and the timestamp of the preceding event is another specific time, the time interval is obtained by calculating the difference between the two.
[0141] Step S412: Divide the operation intensity dimension into two coordinates. The operation force is generated by extracting intensity-related features from the operation vector. The operation frequency within a unit time period is generated by counting the number of times the same operation vector appears within a fixed time window, thus obtaining two sub-coordinate components of the operation intensity dimension.
[0142] The dual-coordinate partitioning of the operational intensity dimension divides it into two sub-coordinate components: operational force and operational frequency per unit time period. Operational force is generated by extracting intensity-related features from the operational vector. For example, the operational vector may contain information such as the magnitude of the operational force and speed; by analyzing and extracting this information, the numerical value of the operational force is obtained.
[0143] The frequency of operations within a unit of time period is calculated by counting the number of times the same type of operation vector appears within a fixed time window. For example, if the time window is set to a time period, the number of times the operation vector "change tires" appears within that time period is counted. Through these operations, two sub-coordinate components of the operation intensity dimension are obtained.
[0144] Step S413: Divide the state change dimension into two coordinates. The immediate state performance after the operation is generated by directly extracting the immediate change features in the state vector. The cumulative state performance is generated by accumulating similar features in the historical state vector, thus obtaining two sub-coordinate components of the state change dimension.
[0145] The dual-coordinate partitioning of the state change dimension divides it into two sub-coordinate components: the immediate state performance after the operation and the cumulative state performance. The immediate state performance after the operation is obtained directly from the immediate change features extracted from the state vector. For example, the state vector may contain an immediate improvement in a certain performance indicator of the vehicle after the operation, such as "driving stability improved to a certain extent".
[0146] The cumulative state performance is calculated by summing up similar features in the historical state vector. If there have been multiple "tire change" operations, each operation will lead to a certain improvement in vehicle driving stability. These improvement values are summed up to obtain the cumulative state performance.
[0147] Step S414: Standardize the sub-coordinate components of each dimension. Through linear transformation, convert the original values such as absolute time markers, time intervals, and operational force into standardized coordinate values within a unified range, while preserving the relative magnitude relationship between the values.
[0148] Standardization transforms the original values of sub-coordinate components in each dimension into standardized coordinate values within a unified interval, facilitating comparison and analysis in a three-dimensional event space. Linear transformation maps the original values to a fixed interval, such as a specific range, through a linear transformation.
[0149] In the specific standardization process, for each sub-coordinate component, the range of its original value is first determined, and then the transformation coefficient is calculated based on this range and the target standardization interval. Using a linear transformation formula, the original values are converted into standardized coordinate values. This preserves the relative magnitudes between the values, ensuring that the characteristics of each dimension are accurately reflected in the three-dimensional event space.
[0150] Step S415: Determine the coordinate scale intervals for each dimension based on the distribution characteristics of historical event data. The scale interval for the time dimension is dynamically adjusted according to the event distribution density, the scale interval for the operation intensity dimension is adjusted according to the distribution characteristics of the force of action, and the scale interval for the state change dimension is adjusted according to the fluctuation range of the state value.
[0151] Determining the coordinate scale interval is to rationally divide the scale in the three-dimensional event space and better represent the values of each dimension. The coordinate scale interval of each dimension is adjusted according to the distribution characteristics of historical event data.
[0152] In terms of the time dimension, the event distribution density may be uneven, with events occurring frequently in some time periods and less frequently in others. Therefore, the scale interval is dynamically adjusted according to the event distribution density. The scale interval is set smaller in time periods with dense events to more accurately represent event time, while the scale interval is set larger in time periods with sparse events.
[0153] The scale interval for the operational intensity dimension is adjusted based on the distribution characteristics of the applied force. If the applied force distribution is relatively concentrated, the scale interval is set to be smaller; if the distribution is relatively dispersed, the scale interval is set to be larger. The scale interval for the state change dimension is adjusted based on the fluctuation range of the state value. For state change dimensions with a large fluctuation range, the scale interval is set to be larger; for state change dimensions with a small fluctuation range, the scale interval is set to be smaller.
[0154] When determining the specific scale intervals, first perform statistical analysis on historical event data and plot the distribution curves of values for each dimension. For the time dimension, count the number of events within different time periods and determine the scale interval adjustment strategy based on the quantity distribution. For the operational intensity dimension, analyze the distribution of the force values and determine the scale interval based on the degree of concentration or dispersion of the distribution. For the state change dimension, calculate the fluctuation range of the state values and determine an appropriate scale interval based on the magnitude of the fluctuations.
[0155] Step S416: Map and associate the standardized coordinate values of each dimension with the scale interval to generate a coordinate mapping table for the three-dimensional event space, and record the correspondence between the original values and the standardized coordinate values.
[0156] Coordinate mapping tables are used to facilitate the search and conversion of raw numerical values and standardized coordinate values in a 3D event space. They map and associate standardized coordinate values of each dimension with scale intervals, establishing a correspondence between raw numerical values and standardized coordinate values.
[0157] For example, for the time dimension, the correspondence between the original values of absolute time markers and time intervals and standardized coordinate values is recorded; for the operation intensity dimension, the correspondence between the original values of the operation force and the operation frequency within a unit time period and standardized coordinate values is recorded; for the state change dimension, the correspondence between the original values of the immediate state performance and the cumulative state performance after the operation and standardized coordinate values is recorded.
[0158] When generating the coordinate mapping table, data structures such as dictionaries or tables can be used to store the correspondences. For each original value in each dimension, the corresponding standardized coordinate value is obtained through the previous standardization process, and these two are stored as a pair of data in the coordinate mapping table. In this way, when a query or transformation is needed, the correspondence can be found quickly and accurately.
[0159] Step S420: Perform spatial mapping preprocessing on the bibasic vectors, extract feature components reflecting the force of action and feature components reflecting the frequency of operation from the operation vector, extract feature components reflecting the instantaneous state and feature components reflecting the cumulative state from the state vector, and perform dimension consistency processing on the extracted feature components to match the vector lengths of each component.
[0160] Spatial mapping preprocessing transforms the bibasic vectors into a form suitable for mapping in a three-dimensional event space. Feature components reflecting the intensity and frequency of actions are extracted from the operation vectors; these components accurately represent information in the dimension of operation intensity. Feature components reflecting the immediate state and the cumulative state are extracted from the state vectors; these components embody information in the dimension of state change.
[0161] Since the extracted feature components may come from different vectors and their dimensions and lengths may differ, dimensionality consistency processing is required. Through feature selection and dimensionality adjustment, the vector lengths of each component are matched to ensure accurate mapping in the three-dimensional event space.
[0162] In one implementation, step S420 may specifically include the following steps S421 to S426: Step S421: Perform vector decomposition on the operation vector, splitting it into a basic operation matrix and an intensity coefficient matrix. The basic operation matrix reflects the type attribute of the operation, and the intensity coefficient matrix reflects the force of the operation. Extract the feature components reflecting the force of the operation from the intensity coefficient matrix.
[0163] Vector decomposition of an operation vector breaks it down into a basic operation matrix and an intensity coefficient matrix. The basic operation matrix contains information about the type of operation, such as "changing a tire" or "vehicle maintenance"; the intensity coefficient matrix reflects the force of the operation, such as the magnitude of the force and the speed of the operation.
[0164] For example, the operation vector is represented in matrix form, and it is split into a basic operation matrix and an intensity coefficient matrix through matrix operations. Feature components reflecting the force of the action are extracted from the intensity coefficient matrix, such as specific elements or specific row vectors in the matrix.
[0165] When performing vector decomposition, appropriate matrix operation rules can be designed based on the meaning and structural characteristics of the elements of the operation vector. For example, if some elements of the operation vector are mainly related to the operation type, while other elements are related to the force of action, they can be separated into the basic operation matrix and the intensity coefficient matrix through matrix partitioning or linear transformation.
[0166] Step S422: Perform frequency statistical analysis on the operation vectors, construct a time window based on the time stamp sequence of the event unit, count the number of times the same type of operation vectors appear within the window, and use the statistical results as the feature component reflecting the operation frequency, which together with the force feature component constitutes the operation intensity feature pair.
[0167] Frequency analysis determines the number of times an operation occurs within a given time period. A time window is constructed based on the time stamp sequence of event units, such as setting the time window to a specific time interval. Within this time window, the frequency of occurrence of similar operation vectors is counted.
[0168] For example, count the number of times the operation vector "change tires" occurs within a certain time period. Use the statistical result as a feature component reflecting the frequency of the operation, and combine it with the force feature component to form an operation intensity feature pair. This feature pair can comprehensively represent the information of the operation intensity dimension.
[0169] When constructing the time window, the start and end times of the window are determined based on the time stamp order of the event units. For each time window, the event units are traversed, and it is determined whether the operation vectors within them are of the same type, and the frequency of their occurrence is counted.
[0170] Step S423: Perform vector decomposition on the state vector, splitting it into an immediate state matrix and a historical state matrix. The immediate state matrix reflects the short-term state changes after the operation, while the historical state matrix reflects the long-term cumulative state changes. Extract the feature components reflecting the cumulative state from the historical state matrix.
[0171] Vector decomposition of the state vector splits the state vector into an immediate state matrix and a historical state matrix. The immediate state matrix contains information on state changes in the short term after an operation, such as an immediate improvement in a certain performance indicator of the vehicle after the operation; the historical state matrix reflects long-term cumulative state changes, such as the overall improvement in vehicle performance after multiple operations.
[0172] For example, the state vector can be represented as a matrix, and then split into an immediate state matrix and a historical state matrix through matrix operations. Feature components reflecting the cumulative state can be extracted from the historical state matrix, such as specific elements or specific row vectors in the matrix.
[0173] In the specific vector decomposition, the state vector elements are divided according to their temporal characteristics and the nature of state changes. Elements reflecting short-term state changes are classified into the immediate state matrix; elements reflecting long-term cumulative state changes are classified into the historical state matrix.
[0174] Step S424: Perform timeliness analysis on the state vector, determine the effective duration of the instantaneous state feature component based on the duration of the state change, merge the values of the instantaneous state feature component that exceed the effective duration into the cumulative state feature component, and update the values of the cumulative state feature component.
[0175] Timeliness analysis considers the impact of the duration of state changes on the state vector. Based on the duration description of the state change, the effective duration of the instantaneous state feature components is determined. For example, an improvement in a vehicle's performance index after an operation may only be effective for a period of time; after this period, its impact will gradually weaken.
[0176] If the value of the instantaneous state feature component exceeds the valid duration, it is merged into the cumulative state feature component, and the value of the cumulative state feature component is updated. This can more accurately reflect the long-term state changes of the vehicle.
[0177] When conducting timeliness analysis, information about duration is first extracted from the state change description to determine the effective duration. For each instantaneous state feature component, its generation time and effective duration are used to determine whether it exceeds the range. If it does, its value is merged into the cumulative state feature component according to certain rules, such as direct addition or proportional weighted addition.
[0178] Step S425: Perform dimension alignment processing on the operation intensity feature pairs and the state change feature pairs. Key dimensions are retained and redundant dimensions are deleted through feature selection, so that the vector lengths of the operation intensity feature pairs and the state change feature pairs are consistent.
[0179] Dimension alignment ensures that the vector lengths of operation intensity feature pairs and state change feature pairs are consistent, enabling accurate mapping in the 3D event space. A feature selection method is used to retain key dimensions and remove redundant dimensions.
[0180] For example, operation intensity feature pairs and state change feature pairs may contain information in multiple dimensions, but some dimensions may not contribute much to the overall feature representation or may contain redundant information. Through analysis and filtering, the dimensions that are most important to the feature representation are selected, ensuring that the vector lengths of the two feature pairs are consistent.
[0181] When performing dimension alignment, the dimensions of the operation intensity feature pairs and state change feature pairs can be evaluated first to determine the importance of each dimension. Feature importance evaluation algorithms, such as those based on correlation analysis or machine learning models, can be used. Then, based on the evaluation results, dimensions with lower importance are removed, while key dimensions are retained to achieve consistency in vector length.
[0182] Step S426: Standardize the dimension-aligned operation intensity feature pairs and state change feature pairs, adjust the numerical range of each feature component, unify the dimensions of different feature components, and combine them into a feature group to be mapped.
[0183] Standardization ensures that the numerical ranges of different feature components are consistent, eliminating the influence of dimensions. Standardization of dimension-aligned operation intensity feature pairs and state change feature pairs can be performed using linear transformation methods.
[0184] The values of each feature component are converted into standardized values within a unified range, such as a specific range. This ensures that different feature components are comparable in the three-dimensional event space. The standardized operation intensity feature pairs and state change feature pairs are combined into a feature group to be mapped, preparing for subsequent spatial mapping.
[0185] In the specific standardization process, for each feature component, the range of its original value is determined, and then a linear transformation is performed based on the target standardization interval. For example, if the original value range of a feature component is from the minimum to the maximum value, and the target standardization interval is from a specific lower limit to a specific upper limit, the original value is converted into a standardized value by calculating the difference ratio.
[0186] Step S430: Embed the preprocessed feature components into the three-dimensional event space. The time dimension coordinates are generated by combining the absolute time marker and the time interval. The operation intensity dimension coordinates are generated by fusing the force feature component and the operation frequency feature component. The state change dimension coordinates are generated by weighted combination of the instantaneous state feature component and the cumulative state feature component to obtain the spatial coordinate points.
[0187] Embedding the preprocessed feature components into the 3D event space involves converting these feature components into coordinate points in the 3D space. The time dimension coordinates are generated through a joint transformation of absolute time markers and time intervals. Based on the previously generated coordinate mapping table, the original values of the absolute time markers and time intervals are converted into standardized coordinate values, and then the time dimension coordinates are generated through a specific combination method.
[0188] The operation intensity dimension coordinates are generated by fusing the force feature component and the operation frequency feature component. A weighted summation method can be used to combine the force feature component and the operation frequency feature component according to different weights to obtain the operation intensity dimension coordinates.
[0189] The state change dimension coordinates are generated by a weighted combination of the instantaneous state feature components and the cumulative state feature components. Different weights are assigned to the instantaneous and cumulative states based on their importance, and the two feature components are then summed using weighted methods to obtain the state change dimension coordinates.
[0190] Finally, spatial coordinates are obtained in the three-dimensional event space, and each coordinate point represents the position of an event unit in the three-dimensional space.
[0191] In the specific coordinate generation process, for the time dimension, the standardized coordinate values of absolute time markers and time intervals are combined according to a certain mathematical relationship, such as linear addition or weighted addition. For the operation intensity dimension, the standardized values of the force characteristic component and the operation frequency characteristic component are weighted and summed according to pre-set weights. For the state change dimension, the standardized values of the instantaneous state characteristic component and the cumulative state characteristic component are also weighted and summed according to weights to obtain the final spatial coordinate points.
[0192] Step S440: Construct a geometric constraint rule base for the vehicle's physical state, including time series rules, operation-state association rules, and state evolution rules. The time series rules determine the threshold range through statistical analysis of historical event time intervals, and the operation-state association rules determine directional constraints through correlation analysis of historical operations and state changes.
[0193] The geometric constraint rule base for vehicle physical state ensures that event units in the three-dimensional event space conform to the logic and rules of vehicle physical state. Time series rules specify the reasonable range of the time sequence and time interval of events, operation-state association rules restrict the direction of association between operation behavior and state change, and state evolution rules regulate the changes in vehicle state with time and operation.
[0194] In one implementation, step S440 may specifically include the following steps S441 to S446: Step S441: Construct time series rules, perform statistical analysis on the time-stamped sequence of historical event data, calculate the distribution characteristics of the time intervals between adjacent events, and determine a reasonable range of time intervals based on the distribution characteristics, which serves as the threshold for the time series rules.
[0195] When constructing time series rules, the first step is to collect time-stamped sequences of a large amount of historical event data. Statistical analysis is then performed on these time-stamped sequences to calculate the distribution characteristics of time intervals between adjacent events, such as the mean, median, and standard deviation. Based on these distribution characteristics, a reasonable range for the time intervals is determined.
[0196] For example, if the time intervals between adjacent events are concentrated around a certain value, this value can be used as the center to determine a range of fluctuations as a reasonable range. This reasonable range can then be used as the threshold for time series rules to determine whether the temporal order of events in the three-dimensional event space is reasonable.
[0197] In specific statistical analysis, statistical software or statistical functions in programming languages can be used to process the time-stamped sequence. The difference between adjacent time stamps is calculated to obtain the time interval sequence. Then, descriptive statistical analysis is performed on this sequence to obtain its distribution characteristics. Based on the characteristics of the distribution, combined with domain knowledge and experience, a reasonable range of time intervals is determined.
[0198] Step S442: Construct operation-state association rules, perform association analysis on historical operation vectors and state vectors, calculate the co-occurrence relationship between operation type and state change direction, incorporate operation-direction pairs with significant co-occurrence relationships into the rules, and restrict specific operation types to correspond to specific state change directions.
[0199] When constructing operation-state association rules, correlation analysis is performed on historical operation vectors and state vectors. Statistical analysis is used to calculate the co-occurrence relationships between different operation types and different state change directions. For example, the number of co-occurrences between the "replacing tires" operation and the "improving driving stability" state change direction is counted.
[0200] Incorporate operation-direction pairs with significant co-occurrence relationships into the rules, restricting specific operation types to correspond to specific state change directions. For example, if statistics show that the "replace tire" operation and the "improve driving stability" state change direction co-occur frequently, then "replace tire - improve driving stability" will be included as a rule in the operation-state association rule base.
[0201] When conducting association analysis, data mining algorithms, such as frequent itemset mining algorithms, can be used to identify co-occurrence patterns between operation types and state change directions. These co-occurrence patterns are then evaluated to determine the significance of the co-occurrence relationships, and significant co-occurrence pairs are used as rules.
[0202] Step S443: Construct state evolution rules, perform trend analysis on the historical state vector sequence, calculate the single change amplitude distribution and cumulative change amplitude distribution of each state attribute, and determine the reasonable range of single change amplitude and cumulative change amplitude based on the distribution characteristics, which serves as the threshold for the state evolution rules.
[0203] When constructing state evolution rules, trend analysis is performed on the historical state vector sequence. The distribution of single-time change amplitude and cumulative change amplitude of each state attribute is calculated, such as the distribution of single-time improvement amplitude and cumulative improvement amplitude after multiple operations for vehicle driving stability.
[0204] Based on these distribution characteristics, reasonable ranges for single-transaction and cumulative change amplitudes are determined and used as thresholds for state evolution rules. For example, if the distribution of single-transaction stability improvement amplitudes is concentrated within a small interval, this interval can be considered a reasonable range for single-transaction amplitudes; similarly, a reasonable range for cumulative change amplitudes is determined based on distribution characteristics.
[0205] When conducting trend analysis, time series analysis methods can be used to model and analyze historical state vector sequences. The magnitude of change in state attributes after each operation is calculated to obtain a sequence of single-operation change magnitudes, which is then statistically analyzed to obtain distribution characteristics. For cumulative change magnitudes, the magnitudes after each operation are summed to obtain a cumulative change magnitude sequence, which is then subjected to statistical analysis. Based on the analysis results, combined with domain knowledge and experience, a reasonable range of change magnitudes is determined.
[0206] Step S444: Integrate time series rules, operation-state association rules, and state evolution rules into a geometric constraint rule library. Each rule includes a rule identifier, applicable operation type, applicable state attribute, threshold range, and conflict handling strategy. The conflict handling strategy defines the priority relationship between rules.
[0207] Time series rules, operation-state association rules, and state evolution rules are integrated into a geometric constraint rule base. Each rule is assigned a unique rule identifier, specifying the applicable operation types and applicable state attributes. Simultaneously, the rule's threshold range, such as time interval thresholds and operation-state association directions, is recorded.
[0208] The conflict resolution strategy defines the priority relationships between rules. When conflicts arise between different rules, they are handled according to the priority relationship. For example, if a time series rule and an operation-state association rule conflict on a certain event unit, the conflict resolution strategy determines which rule takes precedence.
[0209] When integrating rules, data structures such as database tables or lists can be used to store rule information. Each rule is a record containing fields such as rule identifier, applicable operation type, applicable status attribute, threshold range, and conflict handling strategy. During storage, it is crucial to ensure the completeness and accuracy of the rule information to facilitate subsequent rule matching and querying.
[0210] Step S445: Verify the rule base using historical event data, calculate the rule matching pass rate, adjust the parameters of rules with a pass rate that does not meet the standard, update the threshold range or conflict handling strategy, until the rule base's matching pass rate meets the requirements.
[0211] When validating the rule base, event units in historical event data are matched against rules in the geometric constraint rule base. For each event unit, its compliance with rules in the rule base is checked in terms of time series, operation-state association, and state evolution. The rule matching pass rate is obtained by calculating the ratio of the number of successfully matched event units to the total number of event units.
[0212] If the rule matching success rate does not reach the pre-set standard, it indicates that some rules in the rule base are inaccurate or unreasonable. In this case, parameter adjustments to these rules are necessary. For time series rules, the threshold range of the time interval may need to be adjusted; for operation-state association rules, the correlation between operation type and state change direction may need to be reassessed; for state evolution rules, the reasonable ranges for single change magnitude and cumulative change magnitude may need to be updated. Simultaneously, the conflict handling strategy may also need to be adjusted to ensure a more reasonable priority relationship between rules.
[0213] When performing validation and adjustments, automated programs can be used to complete rule matching and pass rate calculation. For rules that fail validation, the reasons for the mismatch are analyzed, and targeted parameter adjustments are made based on the reasons. After each adjustment, validation is performed again until the pass rate of the rule base meets the requirements.
[0214] Step S446: Store the verified geometric constraint rule base as structured data, including rule tables, threshold parameter tables, and conflict handling strategy tables, for rule matching detection of spatial coordinate points in the three-dimensional event space.
[0215] The validated geometric constraint rule base is stored as structured data to facilitate subsequent rule matching and detection. The rule table records the basic information of each rule, including rule identifier, applicable operation type, applicable state attribute, etc.; the threshold parameter table stores the threshold range of the rules, such as time interval threshold, operation-state association direction, etc.; the conflict handling strategy table defines the priority relationship between rules and the conflict handling method.
[0216] This structured data can be stored using a database management system, such as a relational database like MySQL or PostgreSQL. Create appropriate table structures to store the rule table, threshold parameter table, and conflict handling strategy table. During storage, ensure data integrity and consistency to provide reliable data support for subsequent rule matching detection.
[0217] Step S450: Perform rule matching detection on the spatial coordinate points in the three-dimensional event space, compare the values of each dimension of the spatial coordinate points with the threshold range of the geometric constraint rule library, identify contradictory coordinate points that exceed the threshold range, and record the event unit and contradiction type corresponding to the contradiction.
[0218] When performing rule matching detection on spatial coordinate points in the 3D event space, the values of the time dimension, operation intensity dimension, and state change dimension of each spatial coordinate point are compared with the threshold range in the geometric constraint rule base. For the time dimension, it is checked whether the time interval is within the reasonable range specified by the time series rules; for the operation intensity dimension and the state change dimension, it is checked whether the operation-state association conforms to the operation-state association rule, and whether the state evolution conforms to the state evolution rule.
[0219] If the value of a certain dimension of a spatial coordinate point exceeds the threshold range of the corresponding rule in the rule base, the coordinate point is considered a contradictory coordinate point. Record the event unit corresponding to the contradictory coordinate point and the contradiction type, such as temporal sequence contradiction, operation-state association contradiction, or state evolution contradiction, etc.
[0220] When performing rule matching detection, a program can be written using a programming language to traverse all spatial coordinate points in the three-dimensional event space and compare their dimensional values with the threshold ranges in the rule base. For contradictory coordinate points, their relevant information is stored in a log file or database table for subsequent analysis and processing.
[0221] Step S460: Adjust the parameters of the bibasic vectors corresponding to the contradictory coordinate points, call the corresponding correction strategy in the rule base according to the contradiction type, adjust the force feature component of the operation vector or the cumulative state feature component of the state vector, so that the corrected spatial coordinate points meet the threshold range of the geometric constraint rule base, and generate a set of correction vectors to eliminate contradictions.
[0222] Once a contradictory coordinate point is identified, its corresponding basis vector needs to be adjusted. Based on the type of contradiction, the appropriate correction strategy is retrieved from the geometric constraint rule library.
[0223] If there is a contradiction in the time sequence, it may be necessary to adjust the time-related feature components of the operation vector to change the time order of events and make it conform to the time series rules. For example, if an operation event occurs too early or too late, causing the time interval to exceed a reasonable range, the time-related parameters in the operation vector, such as the operation start time or operation duration, can be adjusted appropriately.
[0224] If there is a contradiction in the operation-state association, it is necessary to adjust the strength feature component of the operation vector or the state change feature component of the state vector to ensure that the association between the operation and the state change conforms to the operation-state association rule. For example, if an operation should lead to a specific state change, but the actual state change does not conform to the rule, the strength of the operation vector or the magnitude of the state change in the state vector can be adjusted.
[0225] If there is a contradiction in the state evolution, the cumulative state feature components of the state vector need to be adjusted to make the state evolution conform to the state evolution rules. For example, if the cumulative change in the vehicle state exceeds a reasonable range, the parameters related to the cumulative state in the state vector can be appropriately adjusted to bring it back to a reasonable range of change.
[0226] When adjusting parameters, the corresponding feature components of the bibasic vectors are modified according to the requirements of the correction strategy. After each adjustment, the position of the spatial coordinate point is recalculated, and it is checked whether it meets the threshold range of the geometric constraint rule base. If it still does not meet the requirements, the adjustment continues until the corrected spatial coordinate point meets the rule requirements, ultimately generating a corrected vector set that eliminates the contradictions.
[0227] Step S500: Perform projection transformation on the modified vector set, project the modified vector set onto the preset used car valuation model data field space, and generate structured data units aligned with the used car valuation model data field space. The structured data units are used to construct a standardized used car dataset.
[0228] The purpose of projection transformation is to convert the corrected vector set, after eliminating inconsistencies, into structured data units that match the data field space of the used car valuation model. The used car valuation model data field space defines various data fields used to assess the value of used cars, such as basic vehicle information, performance indicators, and historical maintenance records. Through projection transformation, the information in the corrected vector set is mapped to these data fields, generating standardized structured data units and providing a foundation for constructing a standardized used car dataset.
[0229] In one implementation, step S500 may specifically include the following steps S510-S560: Step S510: Perform hierarchical parsing of the data field definitions in the used car valuation model, extract the parent-child dependency relationship and sibling relationship between fields, and construct a field dependency network containing field hierarchical paths and association weights. The parent-child dependency relationship represents the inclusion relationship between fields, and the sibling relationship represents the collaboration relationship between fields.
[0230] When performing hierarchical parsing of the data field definitions in the used car valuation model, the meaning and function of each data field are first analyzed to determine the hierarchical structure between the fields. Some fields may be subfields of other fields, exhibiting parent-child dependency relationships; other fields are at the same level and have synergistic effects on each other, i.e., sibling relationships.
[0231] For example, the "Vehicle Basic Information" field may contain subfields such as "Vehicle Brand," "Vehicle Model," and "Year of Production," which have a parent-child dependency relationship with the "Vehicle Basic Information" field. Meanwhile, the "Engine Power" and "Torque" fields under the "Vehicle Performance Indicators" field are sibling fields, collectively describing the vehicle's power performance.
[0232] When constructing a field dependency network, a hierarchical path is determined for each field, which is the complete path from the root field to that field. Simultaneously, association weights are assigned to the relationships between fields; a higher association weight indicates a stronger relationship between the two fields. A graph data structure can be used to represent the field dependency network, where nodes represent data fields, edges represent the relationships between fields, and the edge weights represent the association weights.
[0233] When performing hierarchical parsing and network construction, data structures and algorithms from programming languages can be used. Data field definitions are stored in a data structure, such as a dictionary or list, and the field dependency network is constructed by traversing and analyzing this data.
[0234] Step S520: Perform feature layering processing on the modified vector set. Based on the business impact of the vector dimension, divide the two basis vectors into core feature vectors and auxiliary feature vectors. The core feature vectors correspond to the key data fields in the valuation model, and the auxiliary feature vectors correspond to the secondary data fields.
[0235] When performing feature layering on the modified vector set, core feature vectors and auxiliary feature vectors are divided according to the business impact of the vector dimensions. Core feature vectors contain features that have a significant impact on used car valuation; these features correspond to key data fields in the used car valuation model, such as vehicle mileage and engine condition. Auxiliary feature vectors contain features that have a relatively smaller impact on used car valuation; these correspond to secondary data fields, such as the color of the vehicle's interior and some minor exterior decorations.
[0236] In the specific segmentation, the business impact of each vector dimension can be determined by leveraging the experience and knowledge of domain experts and analyzing historical data. Vector dimensions with high business impact are combined into core feature vectors, while vector dimensions with low business impact are combined into auxiliary feature vectors.
[0237] Step S530: Perform multi-path mapping on the core feature vector based on the field dependency network. Map a core feature vector dimension to multiple related fields according to the field hierarchical path. Allocate the feature proportion of each field through the association weight to generate the preliminary field mapping result.
[0238] When performing multi-path mapping on core feature vectors using field dependency networks, for each dimension of the core feature vector, multiple data fields associated with it are found according to the field hierarchy path. Since a core feature vector dimension may be related to multiple fields, the feature proportion of that dimension in each associated field needs to be allocated according to the association weights in the field dependency network.
[0239] For example, the "vehicle dynamic performance" dimension in the core feature vector may be related to multiple fields such as "engine power," "torque," and "acceleration time." Based on the association weights between these fields and "vehicle dynamic performance" in the field dependency network, the feature values of "vehicle dynamic performance" are proportionally distributed to these associated fields.
[0240] In one implementation, step S530 may specifically include the following steps S531 to S536: Step S531: Extract semantic labels from the core feature vector dimension, generate multi-level semantic labels based on the business description of the vector dimension. The semantic labels include operation type, target object and scope of influence. The fields corresponding to each level of label depend on different path nodes in the network.
[0241] When extracting semantic labels from the core feature vector dimensions, the business description of the vector dimension is analyzed to extract key information and generate multi-level semantic labels. For example, for the "vehicle power performance" dimension, its semantic labels may include "power operation type (such as engine performance improvement operation)," "object of action (vehicle engine)," and "scope of influence (overall vehicle power performance)." These semantic labels can help to more accurately understand the meaning of the vector dimension and correspond to the path nodes in the field dependency network.
[0242] When extracting semantic tags, natural language processing techniques, such as lexical analysis and syntactic analysis, can be used to process the business description in the vector dimension, extract key semantic information, and organize it into semantic tags according to a hierarchical structure.
[0243] Step S532: Traverse the hierarchical path of the field dependency network, match the node fields on the hierarchical path of the field according to the semantic label, and record the field paths and path lengths of the successfully matched fields. The path length represents the hierarchical distance between the semantic label and the field node.
[0244] When traversing the hierarchical path of the field dependency network, the extracted semantic tags are matched against the node fields on the field hierarchy path. If the semantic tag matches the meaning of a node field, the match is considered successful. The successfully matched field path and its length are recorded; the path length reflects the hierarchical distance between the semantic tag and the field node in the field dependency network.
[0245] For example, if the semantic label "Power Operation Type - Engine Performance Enhancement Operation" successfully matches a node field on the path "Vehicle Performance Index - Engine Performance - Power Enhancement" in the field dependency network, record the path and its length.
[0246] Step S533: Calculate the association weights between the core feature vector dimension and each matching field. The association weights are negatively correlated with the path length. At the same time, combine the importance of the field in the hierarchy of the dependency network to generate a weight distribution vector.
[0247] When calculating the association weights between the core feature vector dimensions and each matching field, path length and the field's hierarchical importance in the dependency network are considered. Association weights are negatively correlated with path length; that is, the shorter the path, the higher the association weight. Simultaneously, the field's hierarchical importance in the dependency network also affects the association weight; the more important the field, the higher its association weight.
[0248] For example, for a successfully matched field path, a base weight is calculated based on the path length, and then adjusted according to the field's hierarchical importance to generate the final association weight. The association weights of all matched fields are then combined into a weight distribution vector.
[0249] Step S534: Distribute the values of the core feature vector dimension to each associated field according to the weight distribution vector. The distributed value is equal to the product of the core feature vector dimension value and the corresponding weight, generating the preliminary mapping value of each field.
[0250] When distributing the values of the core feature vector dimension to each associated field according to the weight distribution vector, for each matching field, the value of the core feature vector dimension is multiplied by the associated weight corresponding to that field to obtain the initial mapping value of that field.
[0251] For example, if the value of the core feature vector dimension "vehicle power performance" is a certain value, according to the weight distribution vector, this value is proportionally allocated to related fields such as "engine power", "torque", and "acceleration time" to obtain the preliminary mapping values of these fields.
[0252] Step S535: Perform a weighted summation of the multi-path mapping values of the same field. When a field is mapped by multiple core feature vector dimensions, the summation weight is assigned according to the business impact of each dimension, with higher weights for dimensions with higher business impact.
[0253] When a field is mapped to multiple core feature vector dimensions, multiple preliminary mapped values are obtained. These preliminary mapped values are then weighted and summed, with the summation weight assigned according to the business impact of each dimension. Dimensions with higher business impact receive higher weights to ensure that the final mapped value of the field better reflects its actual situation.
[0254] For example, if the "engine power" field is mapped to the two core feature vector dimensions of "vehicle power performance" and "engine optimization operation effect", the initial mapping values of the two dimensions are weighted and summed according to the business impact of these two dimensions to obtain the final mapping value of the "engine power" field.
[0255] Step S536: Check the field value range in the preliminary field mapping result, truncate the mapping values that exceed the reasonable value range of the field to the boundary value of the field value range, and generate a preliminary field mapping result that meets the field value constraints.
[0256] When checking the range of field values in the initial field mapping results, a reasonable range of values is set for each field. If the mapped value of a field exceeds this range, it is truncated to the boundary value of the range.
[0257] For example, if the reasonable range of values for the "engine power" field is a certain interval, and the initial mapped value exceeds that interval, the mapped value is truncated to the upper or lower limit of the interval to ensure that the field value meets the constraints and generate an initial field mapping result that meets the field value constraints.
[0258] Step S540: Perform supplementary mapping processing on the auxiliary feature vector. Based on the sibling relationship, map the dimensions of the auxiliary feature vector to the fields not covered by the core feature vector. Supplement the field values by weighting the feature proportions to improve the field mapping results.
[0259] When performing supplementary mapping on auxiliary feature vectors, the dimensions of the auxiliary feature vectors are mapped to fields not covered by the core feature vectors based on the sibling relationships in the field dependency network. For each auxiliary feature vector dimension, its feature proportion in the relevant fields is determined, and the values of these fields are supplemented by a weighted method based on the feature proportions.
[0260] For example, if the core feature vector does not cover the "vehicle interior color" field, but the auxiliary feature vector has dimensions related to the vehicle interior, the feature proportion can be determined based on the relationship between the dimension and the "vehicle interior color" field. The feature values of the auxiliary feature vector can then be added to the "vehicle interior color" field proportionally to improve the field mapping result.
[0261] Step S550: Perform field logical relationship checks on the improved field mapping results, analyze the numerical inclusion relationship between parent and child fields and the numerical coordination relationship between sibling fields, identify conflicting fields with logical contradictions, and record the conflict type and corresponding feature vector.
[0262] When checking the logical relationships of the refined field mapping results, analyze the numerical inclusion relationships between parent and child fields and the numerical coordination relationships between sibling fields. The numerical values between parent and child fields should satisfy the inclusion relationship; for example, the value of "Vehicle Basic Information - Vehicle Brand" should be included in the overall description of "Vehicle Basic Information". The numerical values between sibling fields should be coordinated; for example, "Vehicle Performance Indicators - Engine Power" and "Vehicle Performance Indicators - Torque" should be consistent in describing the vehicle's power performance.
[0263] If the numerical relationship between fields is found to be illogical, identify the conflicting fields with logical contradictions, record the conflict type (such as mismatched values between parent and child fields, inconsistent values between sibling fields, etc.) and the corresponding feature vector.
[0264] Step S560: Invoke the field mediation strategy according to the conflict type, adjust the proportion of the core feature vector or auxiliary feature vector corresponding to the conflict field, so that the field value satisfies the correlation relationship of the field dependency network, and generate structured data units aligned with the data field space of the used car valuation model.
[0265] When invoking a field mediation strategy based on the conflict type, different mediation methods are applied for different types of conflicts. For conflicts involving mismatched parent and child field values, it may be necessary to adjust the proportions of the core or auxiliary feature vectors in the parent and child fields to ensure that the child field's value is reasonably contained within the parent field's value range. For conflicts involving inconsistent values in sibling fields, it may be necessary to reallocate the proportions of feature vectors in these sibling fields to ensure their values are compatible.
[0266] In one implementation, step S560 may specifically include the following steps S561 to S566: Step S561: Determine the hierarchical storage structure of the structured data unit, including the event metadata area, core field area, auxiliary field area and association area. The event metadata area stores the unique identifier and timestamp of the event. The core field area stores the fields generated by the core feature vector mapping. The auxiliary field area stores the fields generated by the auxiliary feature vector mapping. The association area stores the hierarchical path and association weight between fields.
[0267] When determining the hierarchical storage structure, structured data units are divided into different areas. The event metadata area stores the unique identifier and timestamp of each event, which helps to quickly locate and distinguish different events. The core field area stores fields generated by the core feature vector mapping; these fields are key information for used car valuation. The auxiliary field area stores fields generated by the auxiliary feature vector mapping, providing supplementary information. The relationship area stores the hierarchical paths and relationship weights between fields, recording the relationships between fields to facilitate subsequent data querying and analysis.
[0268] Step S562: Obtain the pre-built field index system. The index includes field name index, event identifier index and hierarchical path index. The field name index supports querying values by field name, the event identifier index supports querying all fields by event identifier, and the hierarchical path index supports querying parent and child fields by field hierarchical path.
[0269] When acquiring a pre-built field index system, the pre-established index structure is used to improve the efficiency of data querying. Field name indexes can quickly locate the value of the corresponding field based on the field name; event identifier indexes can query all fields corresponding to a given event based on the event identifier; and hierarchical path indexes can query the relationship between parent and child fields based on the field hierarchy path.
[0270] Step S563: Timestamp the field values in the core field area and auxiliary field area, record the generation time of each field value and the corresponding feature vector dimension, and construct a field source chain based on the timestamp and dimension information. The source chain includes the vector dimension of the value source, the mapping weight, and a summary of the generation process.
[0271] When timestamping field values, the generation time and corresponding feature vector dimension of each field value are recorded. Based on this timestamp and dimension information, a field origin chain is constructed. The origin chain helps trace the source of field values, understanding which feature vector dimension it originated from, through what mapping weights and generation process.
[0272] Step S564: Analyze the hierarchical path and association weight of the association area, verify the numerical synergy between the core field and the auxiliary field, the ratio of the core field value to the auxiliary field value must be within the range determined by the association weight, and identify abnormal field pairs that exceed the range.
[0273] When analyzing the hierarchical path and association weights in the association relationship area, examine the numerical synergy between core fields and auxiliary fields. Determine a reasonable range for the ratio of core field values to auxiliary field values based on the association weights, and identify abnormal field pairs that exceed this range. For example, if there is an association weight between the core field "Vehicle Power Performance" and the auxiliary field "Vehicle Power Related Decorations," check whether their numerical ratio is within a reasonable range.
[0274] Step S565: Generate metadata descriptions for structured data units. The metadata includes a storage structure summary, field index directory, traceability chain summary, and abnormal field pair records. The metadata is stored in key-value pair format, supporting fast parsing and querying.
[0275] When generating metadata descriptions for structured data units, information such as storage structure summaries, field index directories, source chain summaries, and records with exception field pairs are compiled into metadata. The metadata is stored in key-value pairs for easy and fast parsing and querying. For example, the storage structure summary can be stored as a key-value pair, with the key being "storage structure summary" and the value being a brief description of the hierarchical storage structure.
[0276] Step S566: Integrate the hierarchical storage structure, field index system, field traceability chain, and metadata description to form a binary structured data unit containing a data area and a metadata area. The data area stores field values in a hierarchical structure, and the metadata area stores indexes and metadata descriptions.
[0277] When integrating the hierarchical storage structure, field indexing system, field traceability chain, and metadata description, the field values are stored in the data area according to the hierarchical storage structure, while the indexes and metadata descriptions are stored in the metadata area, forming binary structured data units. This structure facilitates data storage, transmission, and management, and also makes it easier to query and analyze data using indexes and metadata, providing complete and standardized data units for building standardized used car datasets.
[0278] Figure 2This is a schematic diagram of a hardware entity of a dataset construction device provided in an embodiment of the present invention, such as... Figure 2 As shown, the hardware entity of the dataset construction apparatus 1000 includes a processor 1001 and a memory 1002, wherein the memory 1002 stores a computer program that can run on the processor 1001, and the processor 1001 executes the program to implement the steps in the method of any of the above embodiments.
[0279] Memory 1002 stores computer programs that can run on the processor. Memory 1002 is configured to store instructions and applications executable by processor 1001, and may also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) in processor 1001 and various modules of the dataset construction apparatus 1000. It can be implemented using flash memory or random access memory (RAM). When processor 1001 executes the program, it implements the steps of the data governance-based used car dataset construction method described above. Processor 1001 typically controls the overall operation of the dataset construction apparatus 1000.
[0280] This invention provides a computer storage medium storing one or more programs that can be executed by one or more processors to implement the steps of the data governance-based used car dataset construction method as described in any of the above embodiments.
[0281] The above description is merely an embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for constructing a used car dataset based on data governance, characterized in that, The method includes: The original vehicle description text in the used car transaction scenario is obtained, and the original vehicle description text is cleaned and standardized to obtain a set of vehicle description text to be processed. The set of vehicle description text to be processed contains the status description information and operation record information of the vehicle at different stages of use. Event units are extracted from the set of vehicle description texts to be processed. Event units containing operation behaviors and state changes are identified and extracted from the set of vehicle description texts to be processed. The event units are used to characterize the state change process of the vehicle at the corresponding time node. Each event unit is decomposed into a bibasic vector decomposition process, which decomposes the event unit into a bibasic vector consisting of an operation vector and a state vector. The operation vector is used to characterize the force direction of the active intervention applied to the vehicle, and the state vector is used to characterize the offset magnitude of the vehicle attributes under the intervention. The bibasic vectors are embedded into a preset three-dimensional event space. The bibasic vectors in the three-dimensional event space are detected by the geometric constraint model of the vehicle's physical state. Inconsistent event units are identified and corrected, and a set of corrected vectors that eliminates the contradictions is generated. The modified vector set is subjected to projection transformation processing, and the modified vector set is projected onto the preset used car valuation model data field space to generate a structured data unit aligned with the used car valuation model data field space. The structured data unit is used to construct a standardized used car dataset.
2. The method according to claim 1, characterized in that, The step of extracting event units from the set of vehicle description texts to be processed, and identifying and extracting event units containing operational behaviors and state changes contained in the set of vehicle description texts to be processed, includes: The time stamp parsing is performed on the vehicle description text set to be processed to extract the descriptive information representing time nodes in the text. Based on the time node descriptive information, the vehicle description text set to be processed is divided into text segment units with sequential order. Semantic coherence analysis is performed on adjacent text segment units to calculate the semantic overlap between the end of the preceding segment and the beginning of the following segment. Text segments whose semantic overlap does not meet the set standard are supplemented with content. Semantic role recognition is performed on the text fragment units. The action initiator, action target and state change description in each text fragment unit are analyzed. The specific vehicle component or operation execution subject pointed to by the pronoun is determined by combining the contextual reference relationship, and the labeled text content containing semantic role identifier and reference relationship is obtained. Based on the association between the action initiator and the action target in the annotated text content, the basic compositional relationship of vehicle operation behavior is constructed, the temporal continuity characteristics of the action are analyzed, and one-time operation and continuous operation are distinguished. Based on the semantic association analysis results of the relationship between the state change description and the basic composition of the vehicle operation behavior, the state change description content corresponding to each operation behavior is determined. Combined with the causal relationship of vehicle attributes, the main state changes directly caused by the operation behavior and the related state changes indirectly caused by the operation behavior are identified, and a preliminary set of event units containing the correspondence between operation behavior and state change is obtained. Redundancy checks are performed on the initial set of event units. Based on the combination characteristics of operation behavior type, target, and state change amplitude, the redundancy between event units is calculated. Event units with redundancy reaching the set standard are removed to obtain a deduplicated event unit group. Dependency relationships are identified in the deduplicated event unit group, causal and conditional relationships between event units are analyzed, and preceding and following event identification information is added to each event unit to obtain structured event units containing time sequence, semantic roles and dependencies.
3. The method according to claim 2, characterized in that, The process involves semantic role recognition of the text fragment units, analyzing the action initiator, action target, and state change description in each unit, and determining the specific vehicle component or operation execution subject indicated by pronouns based on contextual referential relationships. This yields annotated text content containing semantic role identifiers and referential relationships, including: The text segment unit is segmented into words, and the continuous text sequence is divided into word components with independent semantic functions. The word components are divided into different categories according to part-of-speech features, and each category of words undertakes different semantic functions in the event description. The context semantic coding model is invoked to perform semantic association processing on the word components, analyze the positional relationship of the word components in the text segment and the semantic influence of adjacent words, and generate a word vector sequence containing context-dependent information. The word vector sequence reflects the semantic connotation of the word components in a specific context. The word vector sequence is input into the semantic role recognition process, each word component is classified into roles, and the semantic role classification result corresponding to each word component is output. The classification result reflects the functional positioning of the word component in the event description. Based on the semantic role classification results, the semantic role identifier of each word component is determined, and the semantic role identifier reflects the functional positioning of the word component in the event description; The semantic role identifier is associated with the corresponding word component and stored to establish a correspondence between word component and semantic role identifier, resulting in an associated data group containing word component, part of speech category and semantic role identifier; Cross-validation is performed on the associated data set, comparing it with common semantic role configurations in the domain, and adjustments are made to the deviations in the recognition results to make the associated data set conform to the semantic specifications of the used car domain, thus obtaining the labeled text content.
4. The method according to claim 1, characterized in that, The step of performing bibasic vector decomposition on each event unit, decomposing the event unit into bibasic vectors composed of operation vectors and state vectors, includes: The event units are structured and transformed from natural language event units into structured event descriptions consisting of operation behavior descriptions, object descriptions, operation implementation tools, environmental influencing factors, and state changes. Each component has a clear logical relationship in the event description. Emerging terms are extracted from the description of vehicle operation behavior and state changes. The extracted emerging terms are then added to the vehicle operation behavior vocabulary set and the vehicle state attribute vocabulary set. The term entries and semantic descriptions in the vocabulary sets are then updated. Based on the updated vehicle operation behavior vocabulary set, the operation behavior descriptions and operation implementation tools in the structured event representation are jointly vectorized. Combining the type characteristics of the operation behavior and the functional characteristics of the operation implementation tools, a comprehensive operation vector is generated that integrates operation type, implementation intensity and tool influence. Based on the updated vehicle state attribute vocabulary set, the state changes and environmental influencing factors in the structured event description are jointly vectorized, distinguishing between recoverable and unrecoverable state changes, adding a continuous impact weight to unrecoverable state changes, and generating a comprehensive state vector that includes the magnitude of state attribute changes and the degree of environmental impact. A causal correlation analysis is performed on the integrated operation vector and the integrated state vector. The co-occurrence of operation behavior and state change in historical event data is statistically analyzed to determine the degree of correlation between the two. Vector pairs with a correlation degree greater than a first threshold are selected, and abnormal vector pairs with a correlation degree less than a second threshold are excluded. The integrated operation vector and integrated state vector are subjected to dimensionality optimization processing to retain key information, resulting in a bibasic vector composed of the operation vector and the state vector.
5. The method according to claim 4, characterized in that, The updated vehicle operation behavior vocabulary set is used to jointly vectorize the operation behavior descriptions and operation implementation tools in the structured event representation. Combining the type characteristics of the operation behavior and the functional characteristics of the operation implementation tools, a comprehensive operation vector is generated that integrates operation type, implementation intensity, and tool influence, including: The operational behavior descriptions in the structured event representation are analyzed in principle. Based on the differences in the effects of the operation on the vehicle's physical structure or electronic system, the operational behavior is divided into different types of effects, and the proportion of each category in the operational behavior description is recorded. The operation implementation tools in the structured event description are functionally deconstructed, and the descriptive information of the tool's functional components, operation mode and accuracy level is extracted to establish a mapping relationship between tool function parameters and operation effect; A dynamic operation dictionary is constructed based on the updated vehicle operation behavior vocabulary set. Operation behaviors of different types are mapped to corresponding entries in the dictionary to generate basic vectors representing operation types. Based on the intensity modifiers in the description of the operation behavior and the analysis results of the tool function parameters, the intensity level of the operation is determined, and the intensity level is converted into an intensity coefficient that matches the basic vector dimension. The base vector and the intensity coefficient are fused element by element to obtain an intermediate vector containing the operation type and implementation intensity. The dimension of the intermediate vector is consistent with the number of entries in the dynamic operation dictionary. Based on the mapping relationship between tool function parameters and operation effects, differentiated weights are assigned to each dimension of the intermediate vector, and a comprehensive operation vector is generated by weighted calculation that integrates operation type, implementation intensity and tool impact. The updated vehicle state attribute vocabulary set is used to jointly vectorize the state changes and environmental influencing factors in the structured event representation, distinguishing between recoverable and irreversible state changes. A persistent impact weight is added to irreversible state changes, generating a comprehensive state vector that includes the magnitude of state attribute changes and the degree of environmental impact, including: The state changes in the structured event description are traced by attributes to determine the core vehicle attributes involved in the state changes and to establish the association between the state change description and the core attributes. An impact analysis is performed on the environmental factors in the structured event description. Based on the duration and scope of the environmental factors, the contribution ratio of each factor to the state change is calculated. An attribute feature space is constructed based on the updated vehicle state attribute vocabulary set. The associated core attributes are mapped to the corresponding dimensions in the feature space to generate basic vectors representing state attributes. Based on the description of changes in the state change and the contribution ratio of environmental factors, calculate the change magnitude value of each core attribute, and convert the change magnitude value into an amplitude coefficient that matches the basic vector dimension. The reversibility of state changes is judged, and recoverable and irrecoverable state changes are distinguished. Time decay weights are added to the amplitude coefficients corresponding to irrecoverable state changes to generate weighted amplitude coefficients. The basic vector and the weighted amplitude coefficients are fused element by element to obtain an intermediate vector containing state attributes and change amplitudes. The dimension of the intermediate vector is consistent with the dimension of the attribute feature space. The values of each dimension of the intermediate vector are adjusted according to the contribution ratio of environmental factors to generate a comprehensive state vector.
6. The method according to claim 1, characterized in that, The step of embedding the bibasic vectors into a preset three-dimensional event space, detecting the bibasic vectors in the three-dimensional event space using a geometric constraint model of the vehicle's physical state, identifying and correcting contradictory event units, and generating a corrected vector set to eliminate contradictions includes: The dimensional composition of the three-dimensional event space is determined. The time dimension includes the absolute time mark of the event occurrence and the time interval between adjacent events. The operation intensity dimension includes the force of the operation and the frequency of the operation within a unit time period. The state change dimension includes the immediate state performance after the operation and the cumulative state performance. Each dimension is transformed into standardized coordinate values within a unified range through linear transformation. The bibasic vectors are preprocessed by spatial mapping. Feature components reflecting the force of action and the frequency of operation are extracted from the operation vector. Feature components reflecting the instantaneous state and the cumulative state are extracted from the state vector. The extracted feature components are processed to ensure that the vector lengths of each component are matched. The preprocessed feature components are embedded into the three-dimensional event space. The time dimension coordinates are generated by the joint transformation of absolute time markers and time intervals. The operation intensity dimension coordinates are generated by the fusion calculation of the action force feature components and the operation frequency feature components. The state change dimension coordinates are generated by the weighted combination of the instantaneous state feature components and the cumulative state feature components to obtain the spatial coordinate points. A geometric constraint rule base for the physical state of a vehicle is constructed, which includes time series rules, operation-state association rules, and state evolution rules. The time series rules determine the threshold range through the statistical analysis of historical event time intervals, and the operation-state association rules determine the directional constraints through the correlation analysis of historical operations and state changes. Rule matching detection is performed on spatial coordinate points in the three-dimensional event space. The values of each dimension of the spatial coordinate points are compared with the threshold range of the geometric constraint rule library to identify contradictory coordinate points that exceed the threshold range and record the event unit and contradiction type corresponding to the contradiction. The parameters of the bibasic vectors corresponding to the contradictory coordinate points are adjusted. Based on the contradiction type, the corresponding correction strategy in the rule base is called to adjust the force feature component of the operation vector or the cumulative state feature component of the state vector so that the corrected spatial coordinate points meet the threshold range of the geometric constraint rule base, and a set of correction vectors to eliminate contradictions is generated.
7. The method according to claim 6, characterized in that, The defined dimensions of the three-dimensional event space include: a time dimension comprising the absolute time marker of the event occurrence and the time interval between adjacent events; an operation intensity dimension comprising the force of the operation and the frequency of operations per unit time period; and a state change dimension comprising the immediate state performance after the operation and the cumulative state performance. Each dimension undergoes a linear transformation to convert the original values into standardized coordinate values within a unified interval, including: The time dimension is divided into two coordinates. The absolute time stamp is converted into a time series value by parsing the time description information of the event unit. The time interval between adjacent events is generated by calculating the difference between the time stamps of the current event and the previous event, thus obtaining two sub-coordinate components of the time dimension. The operation intensity dimension is divided into two coordinates. The operation force is generated by extracting intensity-related features from the operation vector. The operation frequency within a unit time period is generated by counting the number of times the same operation vector appears in a fixed time window, thus obtaining two sub-coordinate components of the operation intensity dimension. The state change dimension is divided into two coordinates. The immediate state performance after the operation is directly extracted from the immediate change features in the state vector. The cumulative state performance is calculated by accumulating similar features in the historical state vector, thus obtaining two sub-coordinate components of the state change dimension. The sub-coordinate components of each dimension are standardized. Through linear transformation, the original values such as absolute time markers, time intervals, and operational force are converted into standardized coordinate values within a unified range, while preserving the relative magnitude relationship between the values. The coordinate scale intervals for each dimension are determined based on the distribution characteristics of historical event data. The scale intervals for the time dimension are dynamically adjusted according to the event distribution density, the scale intervals for the operation intensity dimension are adjusted according to the distribution characteristics of the force of action, and the scale intervals for the state change dimension are adjusted according to the fluctuation range of the state value. The standardized coordinate values of each dimension are mapped and associated with the scale intervals to generate a coordinate mapping table for the three-dimensional event space, recording the correspondence between the original values and the standardized coordinate values; The spatial mapping preprocessing of the bibasic vectors involves extracting feature components reflecting the force of action and the frequency of action from the operation vector, and extracting feature components reflecting the instantaneous state and the cumulative state from the state vector. Dimensional consistency processing is then applied to the extracted feature components to ensure that the vector lengths of each component are matched. This includes: The operation vector is decomposed into a basic operation matrix and an intensity coefficient matrix. The basic operation matrix reflects the type attribute of the operation, and the intensity coefficient matrix reflects the force of the operation. Feature components reflecting the force are extracted from the intensity coefficient matrix. Frequency statistical analysis is performed on the operation vectors. A time window is constructed based on the time stamp sequence of the event unit. The number of times the same type of operation vector appears within the window is counted. The statistical results are used as feature components reflecting the operation frequency and form operation intensity feature pairs with the force feature components. The state vector is decomposed into an immediate state matrix and a historical state matrix. The immediate state matrix reflects the short-term state changes after the operation, while the historical state matrix reflects the long-term cumulative state changes. Feature components reflecting the cumulative state are extracted from the historical state matrix. Perform timeliness analysis on the state vector, determine the effective duration of the instantaneous state feature component based on the duration of the state change, merge the values of the instantaneous state feature component that exceed the effective duration into the cumulative state feature component, and update the values of the cumulative state feature component. The operation intensity feature pairs and state change feature pairs are dimension aligned. Key dimensions are retained and redundant dimensions are removed by feature selection, so that the vector lengths of the operation intensity feature pairs and state change feature pairs are consistent. The operation intensity feature pairs and state change feature pairs after dimension alignment are standardized, and the numerical range of each feature component is adjusted to unify the dimensions of different feature components, and then combined into a feature group to be mapped.
8. The method according to claim 1, characterized in that, The step of performing a projection transformation on the modified vector set, projecting the modified vector set onto a preset used car valuation model data field space, and generating structured data units aligned with the used car valuation model data field space includes: The data field definitions of the used car valuation model are parsed hierarchically to extract the parent-child dependency relationship and sibling relationship between fields. A field dependency network containing field hierarchical paths and association weights is constructed. The parent-child dependency relationship represents the inclusion relationship between fields, and the sibling relationship represents the collaboration relationship between fields. The modified vector set is subjected to feature layering processing. Based on the business impact of the vector dimension, the two basis vectors are divided into core feature vectors and auxiliary feature vectors. The core feature vectors correspond to the key data fields in the valuation model, and the auxiliary feature vectors correspond to the secondary data fields. Based on the field dependency network, the core feature vector is mapped to multiple paths. According to the field hierarchical path, a core feature vector dimension is mapped to multiple related fields. The feature proportion of each field is allocated by the association weight to generate a preliminary field mapping result. Supplementary mapping is performed on the auxiliary feature vectors. Based on the sibling relationship, the dimensions of the auxiliary feature vectors are mapped to fields not covered by the core feature vectors. The field values are supplemented by weighted feature proportions to improve the field mapping results. The improved field mapping results are checked for logical relationships. The numerical inclusion relationship between parent and child fields and the numerical coordination relationship between sibling fields are analyzed. Conflicting fields with logical contradictions are identified, and the conflict type and corresponding feature vector are recorded. Based on the conflict type, the field mediation strategy is invoked to adjust the proportion of the core feature vector or auxiliary feature vector corresponding to the conflict field, so that the field value satisfies the correlation relationship of the field dependency network, and generates structured data units that are aligned with the data field space of the used car valuation model.
9. A dataset construction apparatus, comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 8.