Vehicle-mounted diagnostic data acquisition method, vehicle and storage medium
By planning the dimensions and time periods for data collection in the vehicle system and obfuscating privacy data, the problems of blind collection and weak privacy protection in vehicle diagnostic data collection are solved, and efficient and secure acquisition of diagnostic data is achieved.
Patent Information
- Application Number
- CN202610105000.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-26
- Publication Date
- 2026-03-20
AI Technical Summary
Existing technologies for collecting vehicle diagnostic data suffer from problems such as blindness, incompleteness, and weak privacy protection, making it difficult to efficiently and accurately diagnose and resolve vehicle system problems.
By planning the scope (dimensions and time) of data collection before collection and performing privacy filtering after collection, privacy data is identified and obfuscated to generate final diagnostic data.
It has achieved more accurate, efficient, compliant, and secure diagnostic data collection, improving diagnostic efficiency and resource utilization while ensuring privacy compliance requirements.
Smart Images

Figure CN121704437A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent networked vehicles, and more particularly, to a method for acquiring vehicle diagnostic data, a vehicle, and a storage medium. BACKGROUND
[0002] Modern intelligent networked vehicles integrate a large number of controllers, sensors, and complex software vehicle systems, providing users with rich functions, but also making software faults, interaction abnormalities, and other problems in vehicle operation increasingly diverse and difficult to reproduce. Efficient and accurate diagnosis and resolution of these vehicle system problems are of great significance to ensuring driving safety, improving user experience, and optimizing vehicle research and development and after-sales efficiency.
[0003] In related technologies, vehicle system problems are usually diagnosed using a method of continuously recording or reporting vehicle operation logs at fixed intervals, and when a user discovers a problem, the corresponding period of data is retrieved for backtracking analysis. However, the above method has problems of blind and incomplete data collection and weak privacy protection when collecting data, which need to be solved urgently. SUMMARY
[0004] The present application provides a method for acquiring vehicle diagnostic data, a vehicle, and a storage medium. The method pre-plans the collection range (dimension and time) before data collection and performs privacy filtering after collection, solving the problems of blind and incomplete data collection and weak privacy protection in related technology vehicle diagnostics, and achieving the precision, efficiency, and compliance and safety of diagnostic data collection.
[0005] In a first aspect, a method for acquiring vehicle diagnostic data is provided. The method includes: in response to a diagnostic data collection instruction, determining a data collection dimension and a data collection time period; based on the data collection dimension, collecting target diagnostic data according to the data collection time period; identifying privacy data in the target diagnostic data, and performing fuzzification processing on the privacy data to obtain final diagnostic data.
[0006] Through the above technical solution, by intelligently planning the dimension and time period of data collection after responding to the instruction, the indiscriminate full collection mode is abandoned, unnecessary data generation and transmission are reduced from the source, and the diagnostic efficiency and resource utilization are improved; then, by setting the identification and fuzzification of privacy data as a mandatory link for generating final diagnostic data, it is ensured that the diagnostic demand is met while the data meets the privacy compliance requirements from the beginning of output from the vehicle end.
[0007] In some possible implementation manners, in combination with the first aspect, the determining, in response to the diagnostic data collection instruction, of the data collection dimension and the data collection time period comprises: obtaining, based on the diagnostic data collection instruction, an event semantic description and / or a system event identifier corresponding to the problem trigger event; determining a diagnostic value level of the problem trigger event based on the event semantic description and / or the system event identifier; obtaining event feature information of the problem trigger event, and determining a candidate diagnostic path set from a preset diagnostic scene knowledge graph based on the event feature information, wherein the preset diagnostic scene knowledge graph is constructed according to historical diagnostic cases; determining a key data source set and an effective observation time interval corresponding to the candidate diagnostic path set, and adjusting the key data source set and the effective observation time interval based on the diagnostic value level to obtain the data collection dimension and the data collection time period.
[0008] Through the technical solution, the value of the event is evaluated by fusing the user description and the vehicle-mounted system identifier, so that the resources are preferentially invested in high-value problems. Secondly, the knowledge graph constructed based on historical cases is used for reasoning to quickly generate possible fault paths and required key data and observation time, so that the data collection has clear purpose and pertinence. Finally, the collection range and time preliminarily determined are dynamically adjusted based on the diagnostic value level, so that precise matching between diagnostic resource investment and problem importance is realized.
[0009] In some possible implementation manners, in combination with the first aspect and the above implementation manners, the adjusting, based on the diagnostic value level, of the key data source set and the effective observation time interval to obtain the data collection dimension and the data collection time period comprises: determining a current data source expansion coefficient and a current time window expansion coefficient corresponding to the problem trigger event based on the diagnostic value level; adjusting the key data source set according to the current data source expansion coefficient to obtain the data collection dimension; and adjusting the effective observation time interval according to the current time window expansion coefficient to obtain the data collection time period.
[0010] Through the technical solution, the data source expansion coefficient and the time window expansion coefficient are mapped based on the diagnostic value level, so that the vehicle-mounted system can accurately linearly or nonlinearly scale the core data source and the basic observation time preliminarily circled according to the severity and importance of the problem. This not only realizes the refinement and predictability of diagnostic resource configuration, but also systematically optimizes the overall utilization efficiency of computing, storage and network bandwidth of the vehicle-mounted system on the premise of meeting different levels of diagnostic needs.
[0011] In a possible implementation manner of the first aspect, the adjusting the set of key data sources according to the current data source expansion coefficient to obtain the data collection dimension comprises: calculating a collection priority score of each data source in the set of key data sources based on the preset diagnostic scenario knowledge graph; determining a basic data source subset from the set of key data sources based on the collection priority score; searching, from the preset diagnostic scenario knowledge graph, an expansion data source subset that meets a preset association condition with the basic data source subset based on the current data source expansion coefficient; and generating the data collection dimension based on the basic data source subset and the expansion data source subset.
[0012] By the above technical solution, the data sources are sorted in value by the knowledge graph, the high-priority core data is filtered to form a basic subset, and the minimum data requirement for diagnosis is ensured; then, the expansion data sources associated with the core data are intelligently explored and included in the knowledge graph according to the data source expansion coefficient, so that the collection range can be elastically stretched according to the importance of the event. This strategy not only ensures that the core evidence necessary for diagnosis can be obtained in any case, but also automatically expands the data field to capture more extensive associated clues when facing high-value complex problems, thereby achieving a precise balance between data completeness and collection economy and greatly enhancing the diagnosis capability of complex faults.
[0013] In a possible implementation manner of the first aspect, the adjusting the set of key data sources according to the current data source expansion coefficient to obtain the data collection dimension comprises: calculating a collection priority score of each data source in the set of key data sources based on the preset diagnostic scenario knowledge graph; determining a basic data source subset from the set of key data sources based on the collection priority score; searching, from the preset diagnostic scenario knowledge graph, an expansion data source subset that meets a preset association condition with the basic data source subset based on the current data source expansion coefficient; and generating the data collection dimension based on the basic data source subset and the expansion data source subset.
[0014] By the technical solution, firstly, the preliminary data is used to construct a causal network, a main node of a fault is located, and a theoretical time span of forward tracing of a root cause and backward tracing of influence dissipation is quantified; subsequently, with a center of an initial time interval as an anchor point, the two causal distances are combined with a time window expansion coefficient, and are proportionally extended in a bidirectional manner, so that a final collection time period is generated. This method ensures that the time window can completely cover a key causal chain from occurrence of a potential cause of the fault to complete appearance of an influence, while the depth and breadth of analysis can be flexibly adjusted according to the importance of the event, so that the efficiency and accuracy of root cause analysis of a complex and chain fault are significantly improved while a large amount of irrelevant data is avoided.
[0015] With reference to the first aspect and the above implementation manner, in some possible implementation manners, the calculating the weighted reverse root cause distance and the weighted forward influence distance of the main node comprises: performing reverse traversal in the causal influence network with the main node as a traversal starting point until a first preset stop condition is met, to obtain at least one reverse causal path; calculating a decay cumulative weight of each reverse causal path respectively, and selecting a reverse causal path with the highest decay cumulative weight from all the reverse causal paths as a target reverse causal path; and calculating a time difference between a timestamp corresponding to a starting node of the target reverse causal path and a timestamp corresponding to the main node, to obtain the weighted reverse root cause distance.
[0016] By the technical solution, all possible paths in the causal network are reversely searched, and the overall evidence strength of each path is evaluated by using a decay cumulative weight algorithm, so that the vehicle-mounted system can intelligently identify and lock a single causal chain with the most sufficient evidence and the strongest explanation as an analysis benchmark; finally, an absolute time difference between a starting point (an earliest cause) of the optimal chain and an end point (a current fault) is taken as a reverse distance, so that the forward extension of the time window can accurately cover a core cause period that is most valuable for explaining the current fault, noise caused by a large number of weakly related or irrelevant early events is avoided, and an optimal balance between depth and precision in exploring a root cause is achieved.
[0017] With reference to the first aspect and the above implementation manner, in some possible implementation manners, the calculating the weighted reverse root cause distance and the weighted forward influence distance of the main node comprises: performing forward traversal in the causal influence network with the main node as a traversal starting point until a second preset stop condition is met, to obtain a target node; and calculating a time difference between a timestamp corresponding to the target node and a timestamp corresponding to the main node, to obtain the weighted forward influence distance.
[0018] By the technical solution, the absolute time difference between the final node and the starting node reached through traversing in the forward causal relationship in the causal network from the fault master node until the stopping condition preset by the vehicle-mounted system is met is the weighted forward influence distance. This mechanism ensures that the backward extension of the time window can be accurately terminated at the time when the fault influence is actually dissipated or the vehicle-mounted system is restored to normal, thereby avoiding the continuous collection of invalid data after the fault is settled, ensuring the complete capture of the evolution process of the fault consequences, and realizing precise control of the data collection time cost.
[0019] In combination with the first aspect and the above implementation manners, in some possible implementation manners, the identifying the privacy data in the target diagnostic data and performing blurring processing on the privacy data to obtain final diagnostic data includes: performing multi-modal analysis on the target diagnostic data to obtain visual data flow and / or audio data flow and / or text log data; identifying at least one target image region in the visual data flow, and / or identifying at least one target audio segment in the audio data flow, and / or identifying at least one target log entry in the text log data; determining a diagnostic scene type corresponding to a problem triggering event and a privacy policy based on the diagnostic data collection instruction, and determining a current privacy protection intensity level based on the diagnostic scene type and the privacy policy; performing blurring processing on the at least one target image region, and / or the at least one target audio segment, and / or the at least one target log entry based on the current privacy protection intensity level to obtain the final diagnostic data.
[0020] By the technical solution, first, the multi-modal original data is accurately analyzed and identified to locate specific privacy elements; then, according to the specific scene of the diagnostic event and the compliance policy, a differentiated privacy protection intensity level is dynamically determined; finally, the identified privacy elements are processed to different degrees from basic blurring to advanced content replacement according to the level. This process not only realizes accurate identification and minimal processing of privacy data, avoids damage to the diagnostic value of the "one-size-fits-all" approach, but also, through scene-adaptive intensity adjustment, maximizes data availability for different diagnostic purposes while strictly meeting compliance requirements.
[0021] In a second aspect, an acquisition device for vehicle-mounted diagnostic data is provided, which includes: A determination module configured to determine a data collection dimension and a data collection time period in response to a diagnostic data collection instruction; A collection module configured to collect target diagnostic data according to the data collection time period based on the data collection dimension; A processing module configured to identify privacy data in the target diagnostic data and perform blurring processing on the privacy data to obtain final diagnostic data.
[0022] In a third aspect, a vehicle is provided, comprising a controller, the controller comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, the processor executing the program to implement the method for obtaining vehicle diagnostic data according to the above embodiments.
[0023] In a fourth aspect, a computer program product is provided, the computer program product comprising: computer program code to, when executed on a computer, cause the computer to perform the method for obtaining vehicle diagnostic data according to the first aspect or any possible implementation of the first aspect.
[0024] In a fifth aspect, a computer-readable storage medium is provided, the computer-readable storage medium storing computer program code to, when executed on a computer, cause the computer to perform the method for obtaining vehicle diagnostic data according to the first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF DRAWINGS
[0025] Figure 1 A flowchart of the method for obtaining vehicle diagnostic data provided by the embodiments of the present application is shown. Figure 2 A block diagram of the device for obtaining vehicle diagnostic data provided by the embodiments of the present application is shown. Figure 3 A structural diagram of the vehicle of the embodiments of the present application is shown. DETAILED DESCRIPTION
[0026] The technical solutions in the present application will be described in detail below with reference to the drawings. In the description of the embodiments of the present application, unless otherwise specified, " / " represents the meaning of or, for example, A / B can represent A or B: "and / or" in the text only describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which means that there are three cases of A alone, A and B together, and B alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.
[0027] Hereinafter, the terms "first" and "second" are used only for descriptive purposes, and cannot be understood as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more features.
[0028] Figure 1 A schematic flowchart of the method for obtaining vehicle diagnostic data provided by the embodiments of the present application is shown.
[0029] For example,Figure 1 As shown, the method for acquiring the vehicle diagnosis data comprises the following steps: In step S101, in response to a diagnosis data collection instruction, a data collection dimension and a data collection time period are determined.
[0030] It can be understood that the diagnosis data collection instruction is a command or signal triggering the entire data collection process. The instruction can be initiated by the user (such as long-pressing the fault icon) or automatically detected by the vehicle system (such as monitoring process crashes). The data collection dimension refers to the type of data or the collection of data sources, such as screen recording (visual dimension), in-vehicle audio (auditory dimension), vehicle system error log (text dimension), vehicle CAN (Controller Area Network) bus data (control flow dimension), etc., which defines "what to collect". The data collection time period refers to the specific time window of the data to be collected around the time when the problem occurs, which defines "how long to collect". In the embodiments of the present application, the time window backtracking mechanism is used to collect data, for example, data 60 seconds before the problem occurs to 30 seconds after the problem occurs.
[0031] Specifically, after receiving the diagnosis data collection instruction, the vehicle system can first parse the instruction to understand what problem is currently faced (for example, is it navigation lag or Bluetooth connection failure), and based on this understanding, the data collection dimension and the data collection time period of the diagnosis data to be collected next can be determined. This process can change the blind and full data recording mode in the related art of vehicle diagnosis into a targeted, efficient and on-demand collection mode, which avoids data redundancy from the source and lays a solid foundation for subsequent efficient and accurate analysis.
[0032] In step S102, based on the data collection dimension, the target diagnosis data is collected according to the data collection time period.
[0033] It can be understood that the target diagnosis data is the original, unprocessed, multi-modal data set actually collected from the vehicle system under the two constraint conditions of the data collection dimension and the data collection time period determined in step S101.
[0034] Specifically, after determining the data collection dimension and the data collection time period, the vehicle-mounted system can activate multiple corresponding data collection modules in parallel for data collection. For example, if the visual data dimension is included in the data collection dimension, the screen recording service can be started; if the audio data dimension is included, the microphone audio stream capture can be started; at the same time, the log collection service, the vehicle bus listening service, etc. can also work synchronously within the specified data collection time period. The data collected by all modules can be marked with high-precision synchronization timestamps to ensure that the data from different sources can be accurately aligned on the time axis. Thus, this step eliminates the disadvantages of independent and asynchronous collection of different data sources, and realizes the synchronization and contextualization of multi-modal diagnostic data.
[0035] In step S103, the privacy data in the target diagnostic data is identified, and the privacy data is blurred to obtain the final diagnostic data.
[0036] It can be understood that the privacy data refers to sensitive information that can directly or indirectly identify the identity of a specific natural person or reflect personal activities or habits in the target diagnostic data. For example, faces, license plates, navigation destinations, and contact information in visual data; voice conversation content in audio data; vehicle identification codes, precise GPS coordinates, and phone numbers in text logs, etc. The final diagnostic data is a diagnostic data version that has removed or concealed sensitive personal information after blurring and can be directly used for secure transmission, storage, and subsequent analysis.
[0037] Specifically, after collecting the original data (i.e., the target diagnostic data), the vehicle-mounted system does not directly upload it, but can analyze it using image recognition, voice detection, text analysis, etc. to accurately locate various types of privacy data contained therein. Subsequently, the vehicle-mounted system can use blurring algorithms (such as basic code generation and generative content replacement) to desensitize these privacy information to reduce or eliminate their identifiability while preserving their diagnostic value as much as possible. For example, the face area in the image is processed with mosaic, and the mobile phone number in the log is replaced with partial characters. Thus, this step can eliminate the risk of user sensitive information leakage in the transmission, storage, and analysis process at the source, and fundamentally unifies diagnostic utility and privacy security.
[0038] Thus, by intelligently planning the dimension and time period of data collection first after responding to the instruction, the indiscriminate full-volume collection mode is abandoned, unnecessary data generation and transmission are reduced from the source, and the diagnostic efficiency and resource utilization are improved; subsequently, by setting the identification and fuzzification of private data as a mandatory link for generating the final diagnostic data, it is ensured that the data meets the privacy compliance requirements from the beginning of output from the vehicle end while meeting the diagnostic requirements.
[0039] As a possible implementation manner, in some embodiments, in response to the diagnostic data collection instruction, the data collection dimension and the data collection time period are determined, including: based on the diagnostic data collection instruction, the event semantic description and / or the system event identifier corresponding to the problem trigger event are obtained; based on the event semantic description and / or the system event identifier, the diagnostic value level of the problem trigger event is determined; the event feature information of the problem trigger event is obtained, and based on the event feature information, a candidate diagnostic path set is determined from a preset diagnostic scene knowledge graph, wherein the preset diagnostic scene knowledge graph is constructed according to historical diagnostic cases; the key data source set corresponding to the candidate diagnostic path set and the effective observation time interval are determined, and the key data source set and the effective observation time interval are adjusted based on the diagnostic value level to obtain the data collection dimension and the data collection time period.
[0040] It can be understood that the problem trigger event is defined and carried by the diagnostic data collection instruction, that is, the preliminary information about "what problem has occurred" is implanted when the instruction is generated, such as no reaction to clicking somewhere on the screen, and a timestamp and corresponding event ID are recorded. The event semantic description refers to the problem phenomenon described in natural language provided by the user in the case of triggering the diagnostic data collection instruction by the user (such as "I click the 360-degree view icon of the center control screen, but the screen does not react"), or the default description automatically generated by the vehicle-mounted system according to the user's long press icon (such as "the user reports that the 360-degree view function is abnormal"). The system event identifier refers to the standardized error code or log marker automatically generated by the vehicle-mounted system in the case of triggering the diagnostic data collection instruction by the vehicle-mounted system. For example, the error code in the application log is ERR_APP_CRASH, or the diagnostic trouble code on the vehicle CAN bus is DTC U0155. The diagnostic value level refers to the quantitative level or score obtained after comprehensively evaluating the importance, urgency and analysis value of the current problem trigger event, which can determine how many resources should be invested for diagnosis. The preset diagnostic scene knowledge graph is a structured knowledge base constructed based on historical diagnostic cases, which stores fault phenomena, possible causes, related data, time relationships and other entities and their associations in a graphical form, and is used to simulate expert reasoning. The nodes of the graph can include multiple types of nodes: fault mode nodes (representing a specific system fault or abnormal state (for example: "360-degree view application process dead lock", "Bluetooth module connection timeout")); system state nodes (representing the state of a vehicle subsystem or software component (for example: "central control system CPU (Central Process Unit, Central Processor) occupancy rate > 90%", "CAN bus load is too high"); event / action nodes (representing events triggered by the user or the system (for example: "user clicks icon X", "system starts self-checking")); data source nodes (representing a specific data stream or data item that can be collected. This is the node corresponding directly to the "data dimension set" in the claims); and the like. The edges of the graph represent the association between nodes. The candidate diagnostic path set is a series of possible reason chains or fault hypothesis sets leading to the current problem inferred from the preset diagnostic scene knowledge graph, and each path can represent a possible fault propagation logic. The key data source set is the minimum necessary data type set required to verify each path hypothesis in the candidate diagnostic path set. For example, to verify the path "GPU (Graphics Processing Unit, Graphics Processor) rendering timeout causes lag", the key data sources can include GPU usage rate logs and application process state logs.The effective observation time interval is the shortest time range that needs to be covered in theory to capture the causal relationship described by the candidate diagnostic path set, which can be calculated according to the typical time delay between nodes in the preset diagnostic scene knowledge graph.
[0041] Specifically, the vehicle-mounted system can first obtain only the event semantic description corresponding to the problem trigger event, or only the system event identifier corresponding to the problem trigger event, or both the event semantic description and the system event identifier corresponding to the problem trigger event based on the diagnostic data acquisition instruction, depending on the triggering mode of the diagnostic data acquisition instruction. Based on the event semantic description and the system event identifier, the diagnostic value level of the problem trigger event can be evaluated using a pre-trained model, thereby establishing the priority in resource allocation.
[0042] Then, the vehicle-mounted system can further obtain event feature information of the problem trigger event, including but not limited to event type (such as functional abnormality class (such as UI non-response, network connection loss), performance degradation class (such as high delay, low frame rate), etc.), trigger component identifier (for user trigger, directly derived from the context of user interaction, for example, if the user long-presses the "navigation" application icon, the trigger component identifier is com.company.auto.navi (application package name) or NAVIGATION_MAIN_ACTIVITY (activity name); for system trigger, the responsible component can be parsed from system alerts or error logs, for example, a rendering timeout error may be associated with SurfaceFlinger (graphics composition service) or a specific application process ID), and system alert code (such as automobile diagnostic fault code, application error code). Among them, the event type can be quickly routed to the corresponding fault category in the preset diagnostic scene knowledge graph, narrowing the initial search range. The trigger component identifier is the most direct clue for root cause positioning. In the preset diagnostic scene knowledge graph, the component is one of the core nodes, through which the common fault mode of the component, the dependent services and data sources can be directly associated. The system alert code can provide the most objective and technical fault indication. The alert code can be strongly associated with specific fault detection points, hardware modules or software algorithms in the preset diagnostic scene knowledge graph, and is a key basis for verifying fault hypotheses and deducing data requirements.
[0043] Based on this, based on the event feature information, the preset diagnostic scene knowledge graph can be used for reasoning to generate multiple possible fault cause hypotheses, i.e., a candidate diagnostic path set. For each possible fault path (etiology hypothesis), the vehicle-mounted system can deduce from the preset diagnostic scene knowledge graph which key data sources need to be acquired and in which effective observation time interval in order to verify the hypothesis.
[0044] Finally, the in-vehicle system can dynamically adjust the preliminary evidence collection requirements (key data source set and effective observation time interval) according to the diagnostic value level of the current problem triggering event, to obtain the final data collection dimension and data collection time period.
[0045] In addition, the data collection time period can also be set as a fixed time window, such as 60 seconds before and 30 seconds after the fault, without the need for dynamic adjustment according to the actual situation.
[0046] For example, assuming the current scenario is that the user reports "the music player on the center control screen suddenly has no sound", at this time, the in-vehicle system can obtain the semantic description "music player has no sound", and detect the AUDIO_SERVICE_DIED (audio service abnormal termination) identifier in the system log. Using the pre-trained model to make a comprehensive judgment, the problem triggering event affects the core function, and the diagnostic value level is set to "high". In the preset diagnostic scene knowledge graph, starting from "audio output abnormality", several candidate paths are derived: path 1, audio service process crash-no audio output; path 2, Bluetooth audio channel preemption conflict-system audio routing error-no audio output; path 3, system volume management module failure-mute state abnormally locked-no audio output. In order to verify the three paths, the key data sources and effective observation time intervals are determined respectively: path 1 requires audio service process state log and system event log, time interval is 30 seconds before and after the fault; path 2 requires Bluetooth protocol stack log and audio routing strategy log, time interval is 60 seconds before and after the fault; path 3 requires volume control service log and user operation event log, time interval is 45 seconds before and after the fault. Thus, the key data source set can be initially determined to include 6 types of logs: audio service, system event, Bluetooth, audio routing, volume control, and user operation, and the effective observation time interval is 60 seconds before and after the fault. Finally, since the diagnostic value level is "high", the key data source set and the effective observation time interval can be further optimized and adjusted to obtain more accurate data sources, so as to more completely and accurately analyze the problem context.
[0047] Thus, by fusing user description and in-vehicle system identification to evaluate the value of the event, it ensures that resources are preferentially invested in high-value problems; secondly, the knowledge graph constructed by historical cases is used for reasoning to quickly generate possible fault paths and the key data and observation time required, so that data collection has clear purpose and pertinence; finally, based on the diagnostic value level, the preliminary determined collection range and time are dynamically adjusted, realizing the precise matching between diagnostic resource investment and problem importance.
[0048] It should be noted that the pre-trained model mentioned above for evaluating the diagnostic value level of the problem trigger event can adopt the architecture of feature extraction and fusion layer and multi-task evaluation head. The input layer (i.e. the feature extraction layer) of the model can include two channels, one of which can convert system event identifiers (such as DTC (Diagnostic Trouble Code) codes, error log IDs) into dense vectors through an embedding layer. At the same time, meta-features of the problem trigger event, such as event source module (power domain, infotainment domain, etc.), event severity label (such as INFO, WARNING, ERROR, FATAL), are input. The other channel can be used for event semantic description, using a pre-trained lightweight text encoder (such as a small variant of BERT (Bidirectional Encoder Representations from Transformers) or Sentence-BERT) to obtain semantic vector representation of the sentence. At the same time, through a simple sentiment / emergency degree analysis submodule, the sentiment polarity score and keyword emergency degree score in the description text are extracted (for example, the emergency degree is higher for words such as "stuck", "black screen", "unable to"). In the feature fusion layer, the feature vectors output by the two channels can be spliced and deeply fused through a multi-layer perception or attention fusion module.
[0049] In the multi-task evaluation head part, three closely related sub-tasks can be executed in parallel to comprehensively determine the final level. Among them, sub-task 1 can be a classification task (technical severity classification), outputting discrete severity levels (such as: 1 - prompt, 2 - general, 3 - important, 4 - critical). The training labels of this sub-task come from historical data, such as the final repair cost, processing time, whether it involves recall, etc. Sub-task 2 can be a regression task (root cause analysis value regression), outputting a value between 0 and 1, predicting the potential value of expanding the diagnostic knowledge base and discovering new problems for the problem trigger event. The training labels of this sub-task are based on historical cases, such as whether the event is marked as "new failure mode" or "has reproduction value". Sub-task 3 can also be a regression task (user impact estimation), outputting a value between 0 and 1, evaluating the immediate impact of the problem trigger event on user experience and the risk of user complaints it may cause. The training labels of this sub-task can come from user satisfaction scores or complaint levels in historical customer service tickets. Finally, the diagnostic value level can be determined by weighting and integrating the outputs of the three sub-tasks.
[0050] In the process of training the pre-trained model, a training data set can be constructed according to a large number of historical diagnosis cases, each training sample can include input (system event identifier, event semantic description (or summary)), intermediate label (technical severity label, root cause analysis value label, user impact degree label) and final verification label (optional, the label represents the resource level (such as the amount of data collected, the analysis priority) allocated in the actual processing of the case, which can be used as a supervision signal to fine-tune the comprehensive weight (α (weight of subtask 1), β (weight of subtask 2), γ (weight of subtask 3)) of the calculation of the diagnosis value level).
[0051] The training process can be roughly divided into pre-training, multi-task joint training and reinforcement learning fine-tuning. In the pre-training stage, the text encoder part can be pre-trained on a large amount of text using a general language model, and then adapted to the field on the field text (repair report, user feedback). In the multi-task joint training stage, the labeled training data set is input into the model, the model is forward propagated to obtain the output of the three subtasks, the joint loss is calculated using the loss function (such as L_total = λ1*L_tech + λ2*L_rootCause + λ3*L_impact, wherein L_tech is the cross-entropy loss, L_rootCause and L_impact are mean square error losses, and λ1, λ2 and λ3 are weights for balancing the losses of each task), and the model parameters are updated. In the reinforcement learning fine-tuning stage, the trained model can be deployed in a simulation environment or online A / B test; define a reward function, for example, a high-value event is correctly assigned multiple resources and successfully solved to obtain a positive reward; a low-value event is over-allocated resources and wasted to obtain a negative reward. The model can obtain a reward according to the subsequent processing result (simulation or reality) triggered by the value level of its output, and fine-tune the model using methods such as policy gradient to maximize the long-term cumulative reward.
[0052] The trained model can be deployed in a light form on the vehicle side or edge server, and an online learning pipeline is established, that is, after the new diagnosis case closed loop ends (that is, the problem is solved), the real processing process and result of the case can be automatically labeled and added to the training pool for regular updating of the model, so that it can adapt to new fault modes and business strategy changes.
[0053] For the construction of the preset diagnosis scene knowledge graph, the original data source can include but is not limited to a historical diagnosis report library (containing vehicle problem work orders that have been closed, have a clear root cause and a solution), vehicle operation logs and sensor data (complete vehicle-mounted data (CAN (Controller Area Network, controller area network) signals, application logs, DTC codes, etc.) within a period of time before and after the failure associated with the above-mentioned report), repair records and OTA update logs (recorded specific code modification, parameter adjustment or software patch for each problem), expert knowledge documents (maintenance manual, fault tree analysis report, system architecture specification, etc. Semi-structured documents). Further, using natural language processing technology, key entities are automatically extracted from the diagnosis report and log, for example, fault phenomenon (such as "central control black screen", "360-degree view lag", "Bluetooth connection failure"), system component (such as "GPU module", "Bluetooth protocol stack"), data source (such as "CPU temperature log", "SurfaceFlinger rendering log"), root cause (such as "memory leak", "third-party application conflict"), solution (such as "restart the application", "update the driver"). Then, the relationship type between the entities is labeled by a domain expert or through an algorithm, for example, fault phenomenon - appears as - data source exception, system component - depends on - another component, root cause - causes - fault phenomenon, etc. These entities are stored in a graph database as nodes of the graph, the labeled relationships are stored as edges, and an initial weight is assigned to each edge, which can be set based on co-occurrence frequency (the number of times two entities appear simultaneously in historical cases) or expert experience. Thus, an initial diagnosis scene knowledge graph is obtained.
[0054] Next, the information describing the same fact but existing contradictions in the initial diagnosis scene knowledge graph from different sources is checked and unified, conflict resolution and fusion are realized. With the continuous input of new cases, the weight of the edge can also be dynamically updated through an algorithm (such as based on association rule mining, Bayesian network). For example, if "A component memory leak" and "B application lag" are both observed simultaneously in a large number of new cases, the weight of the causing relationship between them can be automatically enhanced. When the graph-based reasoning successfully guides a diagnosis and is verified, the nodes and edge relationships on the path used in this reasoning are positively reinforced (the weight is increased). On the contrary, if the reasoning fails, the corresponding path can be weakened or marked for review.
[0055] Thus, by constructing and continuously optimizing a dynamically evolving diagnosis knowledge system, the embodiments of the application can systematically transform discrete, unstructured historical experience and real-time data streams into a structured knowledge network that can be calculated, reasoned and self-learned.
[0056] As a possible implementation manner, in some embodiments, the key data source set and the effective observation time interval are adjusted based on the diagnostic value level to obtain the data collection dimension and the data collection time period, including: determining the current data source expansion coefficient and the current time window expansion coefficient corresponding to the problem trigger event based on the diagnostic value level; adjusting the key data source set according to the current data source expansion coefficient to obtain the data collection dimension; and adjusting the effective observation time interval according to the current time window expansion coefficient to obtain the data collection time period.
[0057] It can be understood that the data source expansion coefficient and the time window expansion coefficient are both control parameters determined by the diagnostic value level. The data source expansion coefficient does not directly linearly scale the key data source set, but determines the "breadth" or "depth" of subsequent intelligent screening and exploration. The time window expansion coefficient is used to dynamically expand the length of the initially determined effective observation time interval. The higher the coefficient, the longer the time window will be extended forward (to trace back further reasons) and / or backward (to observe longer influences).
[0058] Specifically, after determining the "benchmark" data requirements (i.e., the key data source set) and the observation window (i.e., the effective observation time interval) of the diagnosed problem, the current data source expansion coefficient and the current time window expansion coefficient can be determined based on the diagnostic value level by querying the preset diagnostic value level-data source expansion coefficient-time window expansion coefficient mapping table (as shown in Table 1). When adjusting, the vehicle-mounted system can use the current data source expansion coefficient to adjust the key data source set to cope with the hidden associations that may exist in complex and high-value problems, to obtain the final data collection dimension. At the same time, the effective observation time interval is adjusted using the current time window expansion coefficient, especially to trace back further forward to explore deeper root causes, and to appropriately extend the backward observation to capture more complete subsequent evolution, thereby obtaining the final data collection time period.
[0059]
[0060] Therefore, by mapping the data source expansion coefficient and the time window expansion coefficient through the diagnostic value level, the vehicle-mounted system can accurately linearly or nonlinearly scale the initially identified core data source and basic observation time according to the severity and importance of the problem. This not only realizes the refinement and predictability of diagnostic resource allocation, but also systematically optimizes the overall utilization efficiency of computing, storage and network bandwidth of the vehicle-mounted system under the premise of meeting the diagnostic needs of different levels.
[0061] As one possible implementation, in some embodiments, the key data source set is adjusted according to the current data source expansion coefficient to obtain the data collection dimension, including: calculating the collection priority score of each data source in the key data source set based on a preset diagnostic scenario knowledge graph; determining a basic data source subset from the key data source set based on the collection priority score; searching for extended data source subsets that satisfy preset association conditions with the basic data source subsets from the preset diagnostic scenario knowledge graph based on the current data source expansion coefficient; and generating the data collection dimension based on the basic data source subsets and the extended data source subsets.
[0062] Understandably, the collection priority score is a numerical value calculated for a data source to quantify its relative importance and urgency in this diagnosis. A higher score means the data source is more critical and irreplaceable for verifying the current fault hypothesis. The basic data source subset consists of the highest-priority data sources selected from the key data source set based on their collection priority scores. This is the core evidence set necessary for the diagnosis and must be collected regardless of the problem's value. The extended data source subset consists of other data sources in the pre-defined diagnostic scenario knowledge graph that have some connection to the basic data source subset (e.g., sharing the same component, similar timing, or causal relationship) but are not included in the key data source set. These are potential auxiliary or exploratory evidence. Pre-defined association conditions refer to the rules used to find extended data sources in the pre-defined diagnostic scenario knowledge graph, such as "no more than 2 hops away from the basic data source node," "sharing the same upstream component with the fault event," and "overlapping with the basic data source in timestamps."
[0063] Specifically, firstly, the in-vehicle system can utilize the domain knowledge (such as the contribution of a data source to the diagnosis of a specific fault mode) contained in a pre-defined diagnostic scenario knowledge graph to calculate a collection priority score for each data source in the key data source set. Then, based on this score, a threshold is drawn, and data sources with collection priority scores greater than a pre-defined threshold are selected to form a basic data source subset. This ensures that, under any circumstances, the diagnosis can obtain the most core and critical evidence, establishing a safety baseline for data collection. Subsequently, the in-vehicle system can use the basic data source subset as an anchor point and the current data source expansion coefficient as an "exploration radius" (or "exploration depth") to actively search for extended data source subsets that meet pre-defined association conditions within the pre-defined diagnostic scenario knowledge graph. The larger the current data source expansion coefficient and the more lenient the search conditions (e.g., a larger limit on the number of association hops), the more extended data sources are included, thus achieving an intelligent extension from "necessary evidence" to "association clues." Finally, the union of the basic data source subset that ensures diagnostic effectiveness and the extended data source subset that provides additional analysis dimensions forms the final data collection dimension. This mechanism ensures a balance between the reliability (basic set guarantee) and completeness (extended set supplement) of data collection.
[0064] The specific process for calculating the collection priority score can be as follows: Assume the key data source set is {S1,S2,...,Sm}, and the candidate diagnostic path set is {P1,P2,...,Pn}, where each path contains a series of fault hypothesis nodes. For each data source Sm and each diagnostic path Pn in the key data source set, find the shortest association path between all nodes on the diagnostic path Pn and the data source Sm in the preset diagnostic scenario knowledge graph, and take the product (or minimum) of the weights of all edges on the diagnostic path as the support score Q(m,n) of the data source Sm for the diagnostic path Pn. If no association is possible, the support score is 0. Since there are multiple candidate diagnostic paths for the current fault, a data source may support multiple diagnostic paths. Its overall support Support(m) can be the maximum value of the support scores of all paths (representing that it can strongly support at least one possibility), or a weighted sum (the weights are the current confidence scores of each path Path_Confidence_n), that is: Support(m)=max[Q(m,1),Q(m,2),...Q(m,n)], Or Support(m)=Σ[Path_Confidence_n*Q(m,n)].
[0065] If the data source Sm is a unique monitoring data source for a specific component or module, its scarcity is high and should be given bonus points. Therefore, the scarcity factor R(m) can be set to 1+δ (δ>0). In addition, the cost of collecting data source Sm (such as CPU overhead, data volume, power consumption) can be estimated. The higher the cost, the more points should be deducted. Therefore, the collection cost factor C(m) can be set to 1 / (1+cost).
[0066] At this point, the base collection priority score of data source Sm is: Base_Score(m) = Support(m) * R(m) * C(m).
[0067] Because faults are dynamic and require real-time monitoring, to make the collection priority scores for each data source more targeted, we can examine the typical temporal correlation between data source Sm and the current problem-triggered event in history. If the preset diagnostic scenario knowledge graph records that data source Sm typically experiences anomalies within x seconds before the problem-triggered event, and the preliminary data also shows that data source Sm has anomalies within the time window [Tx, T], then a high score can be given. Therefore, the larger the temporal proximity factor T(m), the higher the score. This can be achieved by calculating the mutual information or correlation strength between the data sequence of data source Sm and the event stamp.
[0068] Simultaneously, a predefined component association table can be queried to match the metadata of the data source Sm (such as the monitored components and parameter types) with the event characteristic information of the current problem-triggered event (such as the triggering component identifier and system alarm code). The higher the matching degree, the greater the score bonus; therefore, the larger the event characteristic matching factor M(m). For example, if the data source entity is equal to the fault core entity (i.e., exact match), then M(m) = 1.0; if the data source entity is a direct dependency, parent module, or child module of the fault core entity (i.e., direct association), then M(m) = 0.7; if associated through an intermediate entity (i.e., indirect management), then M(m) = 0.3; if there is no association, then M(m) = 0.
[0069] Priority_Score(m) = Base_Score(m) * w1 + T(m) * w2 + M(m) * w3, where w1, w2, and w3 are weight coefficients used to balance static support and dynamic context, which can be preset or learned.
[0070] Where T(m) = max(0, 1 - |Δt| / W), |Δt| is the time difference from the start of the anomaly to the time of the failure, and W is the length of the associated time window (based on knowledge graphs or experience, one or more typical associated time windows related to this type of failure can be determined. For example, for "application lag", the GPU utilization rate may be abnormal within 5 seconds before the failure; for "service crash", the memory utilization rate may continue to increase 30 seconds before the crash).
[0071] Therefore, by prioritizing data sources using a knowledge graph, a high-priority core data set is selected to form a basic subset, ensuring the minimum data requirements for diagnosis. Subsequently, based on the data source expansion coefficient, the knowledge graph intelligently explores and incorporates extended data sources associated with the core data, allowing the scope of data collection to flexibly scale according to the importance of the event. This strategy ensures that the core evidence necessary for diagnosis can be obtained under any circumstances, while automatically expanding the data scope to capture broader related clues when facing high-value and complex problems. This achieves a precise balance between data completeness and collection economy, greatly enhancing the diagnostic capabilities for complex faults.
[0072] As one possible implementation, in some embodiments, the effective observation time interval is adjusted according to the current time window expansion coefficient to obtain the data collection period, including: obtaining the initial diagnostic dataset within the effective observation time interval and constructing a causal influence network based on the initial diagnostic dataset; determining the master node corresponding to the problem triggering event from the causal influence network and calculating the weighted reverse root cause distance and weighted positive influence distance of the master node; calculating the start time of the data collection period based on the center time of the effective observation time interval, the weighted reverse root cause distance and the current time window expansion coefficient, and calculating the end time of the data collection period based on the center time of the effective observation time interval, the weighted positive influence distance and the current time window expansion coefficient.
[0073] Understandably, the initial diagnostic dataset refers to the first, unprocessed raw dataset collected within the effective observation period, based on the initially determined data collection dimensions. These initially determined data collection dimensions can be obtained by querying preset rules based on the event characteristic information of the problem-triggered event. The causal influence network is a graph model where nodes represent key system events (e.g., "service A crash") or state inflection points (e.g., "surge in memory usage"), edges represent potential temporal or logical causal or dependency relationships between nodes, and the weights of the edges represent the strength of these relationships (e.g., the correlation calculated using statistical methods). The master node is the node in the causal influence network that directly corresponds to or has the strongest correlation with the problem-triggered event. The weighted backward root cause distance is a quantified time length that represents the time difference between the time of the most likely root cause event (starting from the master node (fault phenomenon) and tracing backward (towards history) along the causal network, and the time of the master node. It answers the question, "How long do we need to look forward to find the most likely cause?" The weighted positive influence distance is also a quantified time length, which can represent the time difference between the point in time when the influence dissipates or a new steady state is established, starting from the master node and tracing forward (towards the future) along the causal network, and the time point of the master node. It can answer the question, "How long do we need to look back to observe the complete influence?"
[0074] Specifically, the vehicle-mounted system can first construct a causal influence network reflecting the unique causal pattern of this fault using the initial diagnostic dataset within the effective observation time interval. This network is generated in real time, capturing the unique interactions between various system elements in the problem-triggered event. After locating the master node corresponding to the problem-triggered event in this causal influence network, the weighted reverse root cause distance and weighted positive influence distance of the master node can be calculated using intelligent algorithms. These two distances represent the theoretically optimal observation range objectively calculated based on the current data. Finally, the vehicle-mounted system can determine the final data collection period based on the center time of the effective observation time interval, the weighted reverse root cause distance, the weighted positive influence distance, and the current time window expansion coefficient. Specifically, the start time of the data collection period = center time of the effective observation time interval - weighted reverse root cause propagation distance * current time window expansion coefficient; the end time of the data collection period = center time of the effective observation time interval + weighted positive influence dissipation distance * current time window expansion coefficient.
[0075] For example, suppose the current scenario is a user report of a "false alarm triggered by the vehicle's automatic emergency braking system." Based on the initial diagnostic dataset within the effective observation time interval, including data from radar, cameras, and decision module logs, a causal influence network is constructed. Nodes in this network can include "false alarm due to sudden deceleration of the vehicle in front," "brief camera obstruction," "sudden drop in confidence level of the fusion algorithm," and "AEB (Autonomous Emergency Braking) trigger command issued." "AEB trigger command issued" is identified as the master node. Traversing backward from the master node, the edge pointing to the node with the highest weight is found to be the one pointing to "sudden drop in confidence level of the fusion algorithm," and then traversing backward to "brief camera obstruction." The time elapsed along this most likely path is calculated, and considering weight decay, the weighted backward root cause distance is found to be 2.5 seconds (i.e., the obstruction occurs approximately 2.5 seconds before the trigger). Traversing forward from the master node, it is found that the system issues a "false alarm cleared" signal approximately 1.0 second after the trigger, and the state recovers. Therefore, the weighted forward influence distance is 1.0 second. Assuming the effective observation time interval is 10 seconds before and after the fault, the center time is T, and the current time window expansion coefficient is 1.8, then the start time of the data acquisition period is T-(2.5 seconds * 1.8) ≈ T-4.5 seconds; the end time of the data acquisition period is T+(1.0 seconds * 1.8) ≈ T+1.8 seconds. Therefore, the final data acquisition period is determined to be [T-4.5 seconds, T+1.8 seconds].
[0076] Therefore, this method first constructs a causal network using preliminary data, identifies the main fault node, and quantifies the theoretical time span for tracing its root cause forward and tracking the dissipation of its impact backward. Then, using the center of the initial time interval as an anchor point, these two causal distances are combined with a time window expansion coefficient, and extended proportionally in both directions to generate the final data collection period. This method ensures that the time window fully covers the key causal chain of the fault, from the occurrence of potential triggers to the full manifestation of its impact, while also flexibly adjusting the depth and breadth of the analysis according to the importance of the event. This significantly improves the efficiency and accuracy of root cause analysis for complex, chain-like faults while avoiding the collection of large amounts of irrelevant data.
[0077] As one possible implementation, in some embodiments, calculating the weighted reverse root cause distance and weighted positive influence distance of the master node includes: taking the master node as the traversal starting point, performing a reverse traversal in the causal influence network until a first preset stopping condition is met, obtaining at least one reverse causal path; calculating the cumulative decay weight of each reverse causal path, selecting the reverse causal path with the highest cumulative decay weight from all reverse causal paths as the target reverse causal path; calculating the time difference between the timestamp corresponding to the starting node of the target reverse causal path and the timestamp corresponding to the master node, to obtain the weighted reverse root cause distance.
[0078] As can be understood, reverse traversal refers to starting from the master node and searching in the opposite direction to the edge direction. In causal influence networks, this means starting from the result (failure phenomenon) and searching backwards for possible causes. The first preset stopping conditions can include: 1) the traversal reaches the maximum backtracking hop count (e.g., 10 hops); 2) the current node has no incoming edges (i.e., no earlier cause points to it); 3) the cumulative evidence strength of the current path has fallen below a certain threshold. A reverse causal path is a sequence of nodes and edges obtained after reverse traversal, representing a possible causal chain hypothesis from a potential early cause to the current faulty master node.
[0079] Specifically, the vehicle-mounted system, triggered by a faulty master node, reverse-engineers all possible causal chains within the causal influence network, generating multiple reverse causal paths. Subsequently, a decaying cumulative weight algorithm is used to quantify and score each path. This algorithm, by introducing a decay factor *r*, reflects the uncertainty penalty for "remote causes," making the scoring more consistent with real-world logic. The system selects the path with the highest score (i.e., the highest decaying cumulative weight) as the target reverse causal path. Finally, the vehicle-mounted system can calculate the time difference between the start and end points (i.e., the master node) of this optimal path (i.e., the target reverse causal path) as the weighted reverse root cause distance.
[0080] For example, suppose the current scenario is that the music player in a car infotainment system suddenly stutters and stops playing music. First, perform a reverse traversal, starting from the main node where music playback stutters, exploring backwards. Assume there are two paths: Path 1: Stuttering - audio decoding thread blocked - file I / O (Input / Output) error - SD (Secure Digital) card read speed drops sharply; Path 2: Stuttering - system audio service unresponsive - high memory usage - memory leak in a background application. Assuming the attenuation factor r is 0.8, the weights of each edge in path 1 are 0.9, 0.7, and 0.8, and the weights of each edge in path 2 are 0.8, 0.6, and 0.5, then the cumulative attenuation weight of path 1 = 0.9 * 0.8. 0 +0.7*0.8 1 +0.8*0.8 2 =1.972, the cumulative weight of decay for path 2 = 0.8 * 0.8 0 +0.6*0.8 1 +0.5*0.8 2 =1.60. Therefore, path 1 is selected as the target reverse causal path. Assuming that the starting point of path 1, "sudden drop in SD card read speed," occurs 3.2 seconds before the lag, then the weighted reverse root cause distance is 3.2 seconds.
[0081] Therefore, by backward searching all possible paths in the causal network and using a decaying cumulative weight algorithm to evaluate the overall evidence strength of each path, the vehicle system can intelligently identify and lock the single causal chain with the most sufficient evidence and the strongest explanatory power as the analysis benchmark. Finally, the absolute time difference between the starting point (earliest cause) and the ending point (current fault) of this optimal chain is used as the backward distance, thereby ensuring that the forward extension of the time window can accurately cover the core cause period that is most valuable for explaining the current fault, avoiding the introduction of noise due to the inclusion of a large number of weakly correlated or irrelevant early events, and achieving the optimal balance between depth and accuracy when exploring root causes.
[0082] As one possible implementation, in some embodiments, calculating the weighted reverse root cause distance and weighted positive influence distance of the master node includes: taking the master node as the starting point of traversal, performing a forward traversal in the causal influence network until a second preset stopping condition is met to obtain the target node; calculating the time difference between the timestamp corresponding to the target node and the timestamp corresponding to the master node to obtain the weighted positive influence distance.
[0083] As can be understood, forward traversal refers to starting from the main node and searching along the forward direction of the edges (i.e., from cause to effect). In a causal network, this means starting from the current failure and tracing its potential subsequent effects or chain reactions. The second preset stopping condition can include 1) traversing to a node marked as "system steady-state recovery" (such as "service restart completed" or "error flag cleared"); 2) reaching the maximum look-ahead hop count; 3) the outgoing edge weights of the current node are all below a certain threshold (indicating weak subsequent effects). The target node is the last node visited during traversal when the second preset stopping condition is met.
[0084] Specifically, the vehicle-mounted system starts from the faulty master node and explores forward through the causal influence network to trace the propagation path of the fault. When the forward traversal meets the second preset stopping condition, especially when it reaches a node marked "system steady-state recovery," it indicates that the abnormal state caused by the fault has ended, the system has returned to normal operation, and the forward exploration process stops. When the traversal stops, the vehicle-mounted system can record the target node visited at that moment and calculate the time difference between the target node and the master node as the weighted forward influence distance.
[0085] For example, suppose the current scenario is that the "engine malfunction indicator light" on the vehicle's dashboard suddenly illuminates (the main node). First, a forward traversal is performed, starting from the "engine malfunction indicator light" main node. Assume there is a path: malfunction indicator light illuminates - engine enters power-limited mode - cruise control function automatically disables - diagnostic system records steady-state fault codes and completes the recording (this node is marked as "system steady-state recovery"). When the traversal reaches the "diagnostic system records steady-state fault codes and completes the recording" node, the second preset stop condition is triggered, and the traversal stops. This node is the target node. Upon verification, the target node completes the recording 4.5 seconds after the malfunction indicator light illuminates; therefore, the weighted forward influence distance is 4.5 seconds.
[0086] Therefore, by starting from the fault master node and traversing along the positive causal relationship in the causal network until the vehicle system's preset stopping condition is met, the absolute time difference between the final node and the starting node is the weighted positive influence distance. This mechanism ensures that the backward extension of the time window accurately terminates at the moment when the fault's impact actually dissipates or the vehicle system returns to normal, thus avoiding the continued collection of invalid data after the fault has subsided. While ensuring the complete capture of the evolution of the fault's consequences, it also achieves precise control over the time cost of data collection.
[0087] As one possible implementation, in some embodiments, identifying privacy data in the target diagnostic data and blurring the privacy data to obtain the final diagnostic data includes: performing multimodal parsing on the target diagnostic data to obtain visual data streams and / or audio data streams and / or text log data; identifying at least one target image region in the visual data stream, and / or identifying at least one target audio segment in the audio data stream, and / or identifying at least one target log entry in the text log data; determining the diagnostic scenario type and privacy policy corresponding to the problem triggering event based on the diagnostic data acquisition instructions, and determining the current privacy protection strength level based on the diagnostic scenario type and privacy policy; and blurring at least one target image region, and / or at least one target audio segment, and / or at least one target log entry based on the current privacy protection strength level to obtain the final diagnostic data.
[0088] It is understandable that diagnostic scenario types refer to categories categorized based on the source and purpose of diagnostic data collection instructions, such as internal R&D debugging, user-initiated problem reporting, systemic quality monitoring, and incident data recording. Different types have different data privacy requirements. A privacy policy is a predefined set of data processing rules bound to the diagnostic scenario type, specifying the permitted processing methods, retention levels, and anonymization strengths for various types of privacy data under different scenarios.
[0089] Specifically, through multimodal analysis and recognition technology, target diagnostic data can be analyzed in a multimodal manner, separating it into independent data streams such as visual, audio, and text. Then, using image recognition, speech detection, and text analysis techniques, specific segments containing sensitive information in each data stream can be accurately located, such as address information in screen recordings, human voice dialogue in audio recordings, and personal identifiers in logs. After locating privacy information, an adaptive current privacy protection level is dynamically determined based on the diagnostic scenario type associated with the diagnostic data collection command (such as system automatic monitoring alarms or user reports) and the corresponding privacy policy.
[0090] Because the diagnostic data acquisition command itself is a structured data packet, it contains key fields for scenario judgment, such as trigger source identifier (clearly distinguishing whether the command comes from the user interface, system monitoring service, or remote operation and maintenance platform), command type / subtype (e.g., user report_functional abnormality, system alarm_security event, R&D command_active capture), and associated context (the command can carry environmental information of the event, such as whether the vehicle is in factory test mode, R&D debugging mode, or normal driving mode). The in-vehicle system also maintains a scenario type definition library, which can map the features parsed from the diagnostic data acquisition command to standardized scenario types. Common scenario types can include: user active reporting, which is triggered by the in-vehicle user through long-pressing an icon, voice command, etc.; automatic system monitoring alarm, which is automatically triggered by the in-vehicle system background service when it detects a performance, safety, or abnormal threshold exceeding the limit; R&D and test debugging, which is actively triggered by engineers through diagnostic tools during the R&D, testing, or factory stages; and statutory accident data recording, which is triggered by safety modules such as collision sensors when a serious accident is detected.
[0091] A privacy policy is a set of rules that specifies how different types of privacy data should be handled in specific scenarios. Each diagnostic scenario type can be pre-associated with one or more baseline privacy policy templates. For example, if the diagnostic scenario type is user-initiated reporting, the associated policy template for this scenario type is "user data minimization and strong anonymization". After selecting a baseline template, the in-vehicle system can dynamically fine-tune the privacy policy based on the specific characteristics of the problem-triggered event, making it more targeted.
[0092] Privacy protection strength levels can include at least a standard level and an enhanced level. For example, when the diagnostic scenario is internal R&D debugging and complies with data anonymization specifications, the standard level is used. When the diagnostic scenario involves remote reporting of user vehicle problems or when the data will be used for third-party analysis, the enhanced level is mandatory.
[0093] After determining the current privacy protection level, differentiated blurring processing can be performed on privacy data of different modalities according to that level to obtain the final diagnostic data. For the standard level: Gaussian blur or pixel mosaic processing can be applied to visual privacy areas; audio privacy segments can be muted (voiced) or bandwidth limited; text privacy fields can undergo partial character replacement (e.g., masking with asterisks) or general hashing (e.g., taking the MD5 hash of the VIN (Vehicle Identification Number)). For the enhanced level: Generative Adversarial Networks can be applied to visual privacy areas for background fusion replacement, making the area visually natural and free of original privacy information; audio privacy segments can be processed using voiceprint stripping and content resynthesis techniques to generate alternative audio that retains non-speech ambient sounds but eliminates specific semantic content; text privacy fields can be completely deleted or irreversibly encrypted using format-preserving encryption, decryptable only by authorized servers.
[0094] Therefore, the process first involves precise analysis and identification of the multimodal raw data to pinpoint specific privacy elements. Then, based on the specific scenario of the diagnostic event and compliance policies, a differentiated level of privacy protection strength is dynamically determined. Finally, based on this level, the identified privacy elements are processed to varying degrees, from basic blurring to advanced content replacement. This process not only achieves accurate identification and minimal processing of privacy data, avoiding the damage to diagnostic value caused by a "one-size-fits-all" approach, but also, through scenario-adaptive strength adjustment, maximizes data usability for different diagnostic purposes while strictly meeting compliance requirements.
[0095] After completing privacy obfuscation processing and generating compliant final diagnostic data, this application embodiment also constructs a complete automated processing and feedback loop to ensure that the problem is resolved efficiently. First, the vehicle system can perform adaptive data compression and transmission on the diagnostic data packet (i.e., the final diagnostic data) based on the current network conditions and the urgency of the problem. For example, it can use a higher compression rate when the network is poor and prioritize transmission in case of emergency failure, in order to optimize bandwidth utilization and ensure that critical data is delivered to the cloud in a timely manner.
[0096] Once the data is uploaded to the cloud server, it enters the AI (Artificial Intelligence) automatic analysis and processing stage. The cloud-based AI system can perform multimodal analysis of the diagnostic data package, fusing and analyzing screen recordings, audio features, and system logs. The system can match this data with a historical problem database: if it is identified as a known, simple software fault (such as an application process freezing), the AI can automatically generate and issue repair instructions (such as restarting the application); if it is determined to be a complex new problem or a hardware-related fault, it is beyond the scope of AI's automatic processing. For complex problems requiring human intervention, the cloud system can automatically create work orders and assign them to the corresponding development teams, entering the automatic solution execution stage. Programmers can perform in-depth analysis based on the front-end anonymized complete contextual data (video recordings, logs) to accurately locate the root cause and develop repair patches. The solution will be securely deployed to the target vehicle in the form of an OTA (Over-the-Air Technology) software update package or through control commands directly issued from the cloud (such as resetting configuration).
[0097] Finally, a closed-loop issue tracking and feedback mechanism ensures a superior user experience. From the moment an issue is reported, users receive real-time status updates via the in-vehicle interface or mobile application, such as "Issue under analysis" or "Repair solution ready, please upgrade." After the upgrade is complete, the in-vehicle system proactively inquires whether the issue has been resolved and sends confirmation to the cloud, thus forming a complete closed loop from discovery to diagnosis to repair to verification. This not only significantly improves the efficiency and transparency of issue resolution but also continuously accumulates valuable diagnostic case data for optimizing future AI diagnostic models.
[0098] It should be noted that the data analysis stage in this application fully utilizes existing mature and advanced general-purpose artificial intelligence technology stacks. Specifically, the cloud system inputs diagnostic data packets, which have undergone intelligent front-end acquisition and privacy processing, into a deep learning-based multimodal fusion analysis model. This model can leverage currently mature AI sub-modules such as visual understanding, speech recognition, and time-series log analysis to collaboratively analyze screen recordings, environmental audio, and system logs to identify fault modes. Simultaneously, the cloud system uses standard AI algorithms such as knowledge graph embedding and sequence pattern mining to match and reason about the current problem against a historical case library. The entire analysis process does not rely on exclusive AI breakthroughs in a specific field, but rather achieves intelligent diagnosis and classification of complex automotive software problems by cleverly integrating and adapting existing, validated AI models and frameworks.
[0099] Figure 2 This is a schematic diagram of the structure of a vehicle control device provided in an embodiment of this application.
[0100] For example, such as Figure 2As shown, the control device of the vehicle may include: a determination module 100, a data acquisition module 200, and a processing module 300.
[0101] Among them, the determination module 100 is used to determine the data acquisition dimension and data acquisition time period in response to the diagnostic data acquisition command; The data acquisition module 200 is used to collect target diagnostic data based on the data acquisition dimension and the data acquisition time period. The processing module 300 is used to identify privacy data in the target diagnostic data and to blur the privacy data to obtain the final diagnostic data.
[0102] Optionally, in one embodiment of this application, the determining module 100 includes: The acquisition unit is used to acquire the event semantic description and / or system event identifier corresponding to the problem-triggered event based on the diagnostic data acquisition command; The first determining unit is used to determine the diagnostic value level of the problem-triggered event based on the event semantic description and / or system event identifier; The second determining unit is used to obtain event feature information of the problem-triggered event, and based on the event feature information, determine a set of candidate diagnostic paths from a preset diagnostic scenario knowledge graph, wherein the preset diagnostic scenario knowledge graph is constructed based on historical diagnostic cases. The processing unit is used to determine the set of key data sources and the effective observation time interval corresponding to the candidate diagnostic path set, and adjust the set of key data sources and the effective observation time interval based on the diagnostic value level to obtain the data collection dimension and data collection time period.
[0103] Optionally, in one embodiment of this application, the processing unit includes: Determine the sub-unit, which is used to determine the current data source expansion coefficient and the current time window expansion coefficient corresponding to the problem triggering event based on the diagnostic value level; The first adjustment subunit is used to adjust the set of key data sources according to the current data source expansion coefficient to obtain the data collection dimensions; The second adjustment subunit is used to adjust the effective observation time interval according to the current time window expansion coefficient to obtain the data acquisition time period.
[0104] Optionally, in one embodiment of this application, the first adjustment subunit is specifically used for: Based on a pre-defined diagnostic scenario knowledge graph, calculate the collection priority score for each data source in the key data source set; Based on the collection priority score, a subset of basic data sources is determined from the set of key data sources; Based on the current data source expansion coefficient, search for extended data source subsets that meet the preset association conditions with the basic data source subset from the preset diagnostic scenario knowledge graph; Data collection dimensions are generated based on a subset of basic data sources and a subset of extended data sources.
[0105] Optionally, in one embodiment of this application, the second adjustment subunit includes: The building blocks are used to obtain the initial diagnostic dataset within the effective observation time interval and to construct a causal influence network based on the initial diagnostic dataset; The first calculation unit is used to determine the master node corresponding to the problem-triggered event from the causal influence network, and to calculate the weighted reverse root cause distance and weighted positive influence distance of the master node; The second calculation unit is used to calculate the start time of the data collection period based on the center time of the effective observation time interval, the weighted inverse root cause distance, and the current time window expansion coefficient, and to calculate the end time of the data collection period based on the center time of the effective observation time interval, the weighted positive influence distance, and the current time window expansion coefficient.
[0106] Optionally, in one embodiment of this application, the first computing element is specifically used for: Starting from the master node, perform a reverse traversal in the causal influence network until the first preset stopping condition is met, and obtain at least one reverse causal path. Calculate the cumulative decay weight of each reverse causal path, and select the reverse causal path with the highest cumulative decay weight from all reverse causal paths as the target reverse causal path. The weighted reverse root cause distance is obtained by calculating the time difference between the timestamp of the starting node and the timestamp of the main node in the reverse causal path of the target.
[0107] Optionally, in one embodiment of this application, the first computing element is specifically used for: Starting from the master node, perform a forward traversal in the causal influence network until the second preset stopping condition is met, and obtain the target node. Calculate the time difference between the timestamp corresponding to the target node and the timestamp corresponding to the master node to obtain the weighted positive influence distance.
[0108] Optionally, in one embodiment of this application, the processing module 300 is specifically used for: Multimodal analysis is performed on the target diagnostic data to obtain visual data streams and / or audio data streams and / or text log data; Identify at least one target image region in a visual data stream, and / or identify at least one target audio segment in an audio data stream, and / or identify at least one target log entry in text log data; Based on the diagnostic data collection instructions, determine the diagnostic scenario type and privacy policy corresponding to the problem triggering event, and determine the current privacy protection strength level based on the diagnostic scenario type and privacy policy; Based on the current privacy protection level, at least one target image region, and / or at least one target audio segment, and / or at least one target log entry are blurred to obtain the final diagnostic data.
[0109] In summary, the vehicle diagnostic data acquisition device according to the embodiments of this application solves the problems of blind, incomplete, and weak privacy protection in vehicle diagnostics in related technologies by pre-planning the collection range (dimensions and time) before data collection and performing privacy filtering after collection, thereby achieving accurate, efficient, compliant, and secure diagnostic data collection.
[0110] Figure 3 This is a schematic diagram of the structure of a vehicle provided in an embodiment of this application.
[0111] It should be understood that the methods described above can be applied to... Figure 3 In the vehicle with the structure shown.
[0112] like Figure 3 As shown, the vehicle includes a controller, which may include a memory 301 and a processor 302. The memory 301 stores executable program code, and the processor 302 is used to call and execute the executable program code to perform the on-board diagnostic data acquisition method provided in this application embodiment.
[0113] Furthermore, the controller also includes a communication interface 303 for communication between the memory 301 and the processor 302.
[0114] This embodiment can divide the vehicle into functional modules based on the above method example. For example, each module can correspond to a separate function module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0115] It should be noted that all relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.
[0116] It should be understood that the vehicle provided in this embodiment is used to perform the above-described method for acquiring vehicle diagnostic data, and therefore can achieve the same effect as the above-described implementation method.
[0117] This embodiment also provides a computer-readable storage medium storing computer program code. When the computer program code is run on a computer, the computer executes the above-described related method steps to implement the method for acquiring vehicle diagnostic data provided in the above embodiment.
[0118] This embodiment also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned related steps to implement the method for acquiring vehicle diagnostic data provided in the above embodiment.
[0119] In this embodiment, the vehicle, computer-readable storage medium, computer program product, or chip are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding methods provided above, and will not be repeated here.
[0120] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0121] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0122] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for acquiring vehicle-mounted diagnostic data, characterized in that, Includes the following steps: In response to diagnostic data acquisition commands, determine the data acquisition dimensions and data acquisition time periods; Based on the data collection dimensions, target diagnostic data is collected according to the data collection time period. Privacy data is identified in the target diagnostic data, and the privacy data is obfuscated to obtain the final diagnostic data.
2. The method according to claim 1, characterized in that, The step of responding to a diagnostic data acquisition command and determining the data acquisition dimensions and time period includes: Based on the diagnostic data acquisition instructions, obtain the event semantic description and / or system event identifier corresponding to the problem triggering event; Based on the event semantic description and / or the system event identifier, determine the diagnostic value level of the problem-triggered event; Obtain event feature information of the problem-triggered event, and determine a set of candidate diagnostic paths from a preset diagnostic scenario knowledge graph based on the event feature information, wherein the preset diagnostic scenario knowledge graph is constructed based on historical diagnostic cases; The key data source set and effective observation time interval corresponding to the candidate diagnostic path set are determined, and the key data source set and effective observation time interval are adjusted based on the diagnostic value level to obtain the data collection dimension and the data collection time period.
3. The method according to claim 2, characterized in that, The adjustment of the key data source set and the effective observation time interval based on the diagnostic value level to obtain the data collection dimension and the data collection time period includes: Based on the diagnostic value level, determine the current data source expansion coefficient and the current time window expansion coefficient corresponding to the problem triggering event; The key data source set is adjusted according to the current data source expansion coefficient to obtain the data collection dimension; The effective observation time interval is adjusted according to the current time window expansion coefficient to obtain the data acquisition time period.
4. The method according to claim 3, characterized in that, The step of adjusting the set of key data sources according to the current data source expansion coefficient to obtain the data collection dimension includes: Based on the preset diagnostic scenario knowledge graph, calculate the collection priority score of each data source in the key data source set; Based on the collection priority score, a subset of basic data sources is determined from the set of key data sources; Based on the current data source expansion coefficient, search the preset diagnostic scenario knowledge graph for an extended data source subset that satisfies the preset association conditions with the basic data source subset; The data collection dimensions are generated based on the subset of basic data sources and the subset of extended data sources.
5. The method according to claim 3, characterized in that, The step of adjusting the effective observation time interval according to the current time window expansion coefficient to obtain the data acquisition time period includes: Obtain the initial diagnostic dataset within the effective observation time interval, and construct a causal influence network based on the initial diagnostic dataset; Identify the master node corresponding to the problem-triggered event from the causal influence network, and calculate the weighted reverse root cause distance and weighted positive influence distance of the master node; Based on the center time of the effective observation time interval, the weighted inverse root cause distance, and the current time window expansion coefficient, the start time of the data acquisition period is calculated, and based on the center time of the effective observation time interval, the weighted positive influence distance, and the current time window expansion coefficient, the end time of the data acquisition period is calculated.
6. The method according to claim 5, characterized in that, The calculation of the weighted reverse root cause distance and weighted positive influence distance of the master node includes: Starting from the master node, a reverse traversal is performed in the causal influence network until the first preset stopping condition is met, resulting in at least one reverse causal path. Calculate the cumulative attenuation weight for each of the reverse causal paths, and select the reverse causal path with the highest cumulative attenuation weight from all the reverse causal paths as the target reverse causal path; The weighted reverse root cause distance is obtained by calculating the time difference between the timestamp corresponding to the starting node of the target reverse causal path and the timestamp corresponding to the main node.
7. The method according to claim 5 or 6, characterized in that, The calculation of the weighted reverse root cause distance and weighted positive influence distance of the master node includes: Starting from the master node, a forward traversal is performed in the causal influence network until the second preset stopping condition is met, and the target node is obtained. The time difference between the timestamp corresponding to the target node and the timestamp corresponding to the master node is calculated to obtain the weighted positive influence distance.
8. The method according to claim 1, characterized in that, The process of identifying privacy data in the target diagnostic data and obfuscating the privacy data to obtain the final diagnostic data includes: Multimodal analysis is performed on the target diagnostic data to obtain visual data stream and / or audio data stream and / or text log data; Identify at least one target image region in the visual data stream, and / or identify at least one target audio segment in the audio data stream, and / or identify at least one target log entry in the text log data; Based on the diagnostic data collection instructions, determine the diagnostic scenario type and privacy policy corresponding to the problem triggering event, and determine the current privacy protection strength level based on the diagnostic scenario type and the privacy policy; Based on the current privacy protection strength level, the at least one target image region, and / or the at least one target audio segment, and / or the at least one target log entry are blurred to obtain the final diagnostic data.
9. A vehicle comprising a controller, the controller including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the program to implement the method for acquiring vehicle diagnostic data as described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed, implements the method for acquiring vehicle diagnostic data as described in any one of claims 1-8.
Citation Information
Patent Citations
Automobile data analysis method and system based on intelligent diagnostic instrument
CN119541080A
Diagnosis method and device, electronic equipment and storage medium
CN120215468A
Metro equipment fault intelligent diagnosis method and system assisted by large language model
CN120337106A
Mobile application function exception processing method and device, electronic equipment and storage medium
CN120832299A
Vehicle fault root cause diagnosis method, device, equipment and medium
CN121144719A