A computer system installed on the carrier that performs at least one service that is critical to the operational safety of the carrier.

The described system addresses the complexity and cost challenges of ensuring operational safety in on-board systems by using networked redundant computers and cryptographic signatures to detect and manage failures, achieving efficient and cost-effective operational safety.

JP7689524B2Active Publication Date: 2025-06-06COMMISSARIAT A LENERGIE ATOMIQUE ET AUX ENERGIES ALTERNATIVES
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2022530910
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-12-06
Filing Date
2020-11-12
Publication Date
2025-06-06
Estimated Expiration
2040-11-12

AI Technical Summary

Technical Problem

Existing on-board systems for operational safety, such as airbag deployment and automatic emergency braking, face challenges in ensuring safety and availability due to complexity and high costs, particularly when traditional lockstep redundancy is limited to physically close execution units.

Method used

A computer system installed on a carrier communicates with a data concentrator and a monitor in a network, implementing critical safety services redundantly across different computers. This system uses a relay server to calculate and transmit signatures of service execution via a hash chain, allowing for detection of temporal and operational failures, and enabling fail-soft modes and redundancy management.

Benefits of technology

The solution provides operational safety at reduced complexity and cost, enabling effective detection and management of failures, and ensuring high availability of critical services by using networked remote computers and cryptographic signatures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007689524000050
    Figure 0007689524000050
  • Figure 0007689524000051
    Figure 0007689524000051
  • Figure 0007689524000052
    Figure 0007689524000052
Patent Text Reader

Abstract

A computer system (1) installed on a carrier, communicating with a data concentrator (2) and a monitor (M) in the network, and implementing at least one service critical to the operational safety of the carrier, the critical service being implemented by each of the different computers (C1,...,C2) connected to the network. m ) on at least two instances (δ1,...,δ m ) and each computer (C k ) is the critical service instance (δ k ) and configured to perform a critical service under time control.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an on-board hardware and / or software architecture comprised of software services (or components) that interact with devices that concentrate data (e.g., a data concentrator, or "broker," or "oriented services" architecture with a "data function bus"). [Background technology]

[0002] Some functions that are critical to operational safety (such as the deployment of airbags or automatic emergency braking of a vehicle) are performed by complex on-board systems implemented in hardware and software form.

[0003] Developing and preparing (or even ensuring) these systems for proper operational safety can be particularly difficult and expensive, even more so when the underlying hardware or software architecture is complex.

[0004] To improve the safety and availability of critical functions implemented in software, it is customary to deploy multiple redundant implementations in a mode called "lockstep", which involves using at least two physical execution units that run exactly the same code at the same time, e.g. two computers or CPUs (acronym for Central Processing Unit) on the same System-on-Chip (acronym SoC).

[0005] A hardware device built into the system-on-chip detects when the registers (or memory accesses) of the two cores are different, signaling that a fault has occurred in at least one of the two cores and therefore that it needs to switch to a fail-soft mode (e.g., resume functionality, deactivate functionality, signal a fault, engage a secondary execution mode, etc.).

[0006] However, traditional lockstep is only possible between two computers that are physically close to each other to compare their registers or their memories. The present invention allows the use of networked remote computers.

[0007] Traditional lockstep involves the same execution units executing the exact same set of software tasks (so that the execution is the same cycle-by-cycle except in the presence of a fault). Summary of the Invention [Problem to be solved by the invention]

[0008] The object of the present invention is to overcome the aforementioned problems and in particular to provide operational safety at low cost and reduced complexity. [Means for solving the problem]

[0009] One proposal according to one aspect of the invention is a computer system installed on a carrier, communicating with a data concentrator and a monitor in a network and implementing at least one service that is critical to the operational safety of the carrier, the critical service being redundant in at least two instances on different respective computers connected to said network; - each computer implements at least one software task that implements an instance of a critical service; and - an increasing set of task activation dates and a set of corresponding task latest end dates associated with a system initiation, where a threshold difference between the end dates and the corresponding activation dates corresponds to an estimate of a task's execution time or response time; - backup of the internal state of the computer between two successive activations of a service by modeling it by recording the memory state of the task; - updating the internal state of the computer for each activation of the service, starting after the corresponding activation date, by reading input data of the service and calculating output data of the service and providing it to the data concentrator, whereby firstly, the dependency between the updated internal state and the calculated output data and, secondly, the dependency between the previous internal state and the read input data are represented by a transfer function; and - a relay server or proxy configured to calculate a signature characteristic of the execution of an instance of the service from the initial activation of the system to the current latest termination date by a hash chain dependent on a hash function on nb bits, and to transmit the signature to the monitor; configured to perform a critical service under time control by using the A monitor is a computer system that detects failures by analyzing the signatures of instances.

[0010] Such a system makes it possible to provide operational safety at low cost and reduced complexity.

[0011] In one embodiment, the relay server recursively calculates for each instance k within each period (or time step) n by hash chaining using a cryptographic hash function H, such that:

number

number

number

number

number

[0012] Therefore, the value

number

[0013] According to one embodiment, the monitor is configured to detect a temporal failure if the signature for the instance is not received before the latest end date of the current period.

[0014] Thus, the monitor contributes to detecting temporal failures of instances of critical services.

[0015] In one embodiment, the monitor is configured to compare signatures received from the relay server to detect operational failures when the signature of one of the instances differs from the other signatures.

[0016] Thus, the monitor detects whether at least one of the instances of the critical service exhibits an operational impairment.

[0017] In one embodiment, the monitor detects a failure when the signatures of the instances are equal but the internal state and / or output data of one differs from the other.

[0018] Thus, if the number of instances of a service is at least equal to three and less than half of these instances are faulty, the monitor is configured to take a majority vote among the signatures received within a time period and provide a majority signature indicating a functional instance that is different from the majority signature indicating a faulty instance, and is configured to signal the functional and faulty instances to the rest of the system such that transmission of the faulty instances is interrupted on the data concentrator.

[0019] The device thus provides replicated operational and temporal fault detection and isolation as well as outage tolerance, improving the overall availability of the service.

[0020] For example, if an instance d If a fault is detected within a period n, the computer hosting the faulty instance will delay its schedule, possibly with respect to the corresponding latest end date, by applying a transfer function. d and resuming the failed instance from the correct memory state, and for a period n d The monitor is configured to feed back input data from up to the current period to the failed instance, and the monitor is configured to report the failed instance as operational again when the failed instance catches up to the operational instance, i.e., the transfer function has been applied to the input up to period n before the nth latest end date and the signatures are equal again.

[0021] The set of signatures therefore allows identification of failed instances, identification of operational instances (based on which the failed instance is restarted and regenerated), and identification of when the regeneration is completed and the instance becomes operational again.

[0022] In one embodiment, the relay server is a software server implemented on a corresponding computer.

[0023] Thus, the implementation is simplified.

[0024] In one embodiment, the relay server is a hardware server implemented in the data concentrator.

[0025] Thus, proxies are not exposed to the same risks of software executive failure as the instances they monitor.

[0026] According to one embodiment, the system includes a network that is independent of the data concentrator for transmitting signatures by relay servers, the network having a lower passband and higher reliability than that of the data concentrator.

[0027] There is therefore a small risk that the signature comparison will be compromised by integrity defects or time disturbances linked to the transmission between the relay server and the monitor.

[0028] Another proposal according to another aspect of the invention is a method for managing at least one service critical to the operational safety of a computer system installed on a carrier, the critical service being redundant in at least two instances on different respective computers connected to said network, Each implementation of an Instance of a Critical Service shall: - an increasing set of task activation dates and a set of corresponding task latest end dates associated with a system initiation, where a threshold difference between the end dates and the corresponding activation dates corresponds to an estimate of a task's execution time or response time; - backup of the internal state of the computer between two successive activations of a service by modeling it by recording the memory state of the task; - updating the internal state of the computer for each activation of the service, starting after the corresponding activation date, reads input data of the service and calculates output data of the service and provides it to the data concentrator, whereby firstly, the dependency between the updated internal state and the calculated output data and, secondly, the dependency between the previous internal state and the read input data are represented by a transfer function; and - calculation by the relay server of a signature characteristic of the execution of the instance of the service from the initial activation of the system to the current latest termination date, by a hash chain depending on a hash function on nb bits, and transmission of the signature to the monitor; Use; The method for detecting failures by the monitor is by analyzing the signatures of instances.

[0029] The invention will be better understood on studying some embodiments thereof, explained by way of non-limiting examples and illustrated by the accompanying drawings, in which: [Brief description of the drawings]

[0030] [Figure 1a] 1 illustrates a schematic diagram of a system according to an aspect of the present invention; [Figure 1b] 1 illustrates a schematic diagram of a system according to an aspect of the present invention; [Diagram 2] 2 illustrates a schematic of the operation of the system of FIG. 1 including two instances of a service. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0031] Throughout the accompanying drawings, elements having the same reference numbers are similar.

[0032] 1a and 1b show diagrammatically a computer system 1 mounted on a carrier according to two embodiments of the invention.

[0033] A computer system 1 installed on a carrier communicates in a network with a data concentrator 2 and a monitor M and implements at least one service that is critical to the operational safety of the carrier (or is safety-critical). The critical service is redundant, i.e. it is implemented by different respective computers C connected to the network. 1 ,...,C m At least two instances of δ on 1 ,...,δ m (in this case two copies on two respective computers).

[0034] Each computer C 1 ,...,C m is an instance of a critical service, δ k and configured to perform critical services by using: - A series of increasing activation tasks associated with the initiation of the system n and the latest end date of a series of corresponding tasks D n where a threshold gap between the end date and the corresponding activation date represents an increasing set of activation task dates R n and the latest end date of a series of corresponding tasks D n Date R n and D. n conforms to the following inequality:∀n,0 <R n <R n+1 ,0 <D n <D n+1 and D. n -R n ≥ WCET; - The internal state of the computer between two successive activations of the service, modeled by recording the memory state of the computer n (backing up memory state, registers, and variable modeling of code); - Corresponding activation date R n The internal state of the computer for each activation n of the service that starts at n+1The update of the service's input data i n Read the output data of the service n and provide it to the data concentrator 2, firstly, by calculating the updated internal state and the calculated output data s n+1 ,o n and secondly, the dependency between the previous internal state and the read input data s n ,i n The dependence between and is expressed by a transfer function f; - System initial activation 0 to latest end date D n Instances of the service up to δ k Signatures that characterize the execution of

number

number

[0035] A monitor M has an instance δ 1 ,...,δ m Signature

number

[0036] In Figure 1a, the relay server SR k corresponds to the computer C k In Figure 1b, the relay server SR k is a hardware server implemented in the data concentrator 2.

[0037] The hash function H is a cryptographic hash function, i.e., a cryptographic hash function provides a fingerprint h (said to be "fast to compute") that, for messages of any size, is resistant to preimage attacks (given fingerprint h, it is practically impossible to construct a message m such that H(m)=h), to second preimage attacks (it is practically impossible to construct a message m2 such that H(m2)=H(m1), even knowing m1), and to collisions (it is practically impossible to construct two distinct messages m1, m2 such that H(m1)=H(m2).

[0038] If a critical service is redundant, multiple instances δ k (k=1,2...m) implement the same transfer function f but are vulnerable to failures. The variables that model the behavior of instance k are X k and the one that describes an instance without theoretical flaws is called X.

[0039] For each instance δ k is the same time constraint R n , D n and the same initial state

number

number

number

[0040] The instances do not necessarily run simultaneously; they may run on computers with different frequencies or may be blocked by other tasks. The only necessary assumption is that the nth execution or the nth job is activated on the R n and the latest D n The objective is to ensure that the measures are carried out effectively between the

[0041] Instance δ activated during the nth job 1 Assume that has an error: internal fault

number

number

[0042] The present invention is based on the latest date D n For each instance δ k Signatures that characterize the execution of

number

[0043] signature

number

number

number

number

[0044] The system is up to date. n-1 Assuming that the load remains in nominal mode until - signature

number

number

number

number

number

number

number

[0045] False negatives can occur when: - The monitor M itself is in a fault state; this risk is related to at least two monitors M, M ’ can be reduced by various means including monitor redundancy so that - an instance of the service δ k are in a bad state and generate the same signature: - This risk has traditionally been considered sufficiently unlikely to be tolerable in the case of failures of a random and independent or transient nature (if each instance has a probability p of facing a random failure, and if the failures of the instances are considered independent, then the probability that K replicas are in a failed state is p K ), - The risk of permanent or common mode interference has traditionally been reduced by rigorous design, analysis and testing processes; - The signatures are all equal, but the duplicates deviate from the specification:

number

number

number

number

number

[0046] If the monitor M detects a deviation in the received signature, it can signal it to the operational status management or health management device in charge of deactivating the replicas, switching to a fail-soft mode called FT (acronym for Fault Tolerant Mode) or restarting all replicas in a reference state.

[0047] In addition, if more than two copies of a service are instantiated, the monitor M can determine which instances are in a failed state by majority vote and selectively deactivate or resume them.

[0048] As shown in Fig. 2, the fault state instance δ k To resume, we need to use a non-faulty instance, δ j is the final internal state of the nth iteration.

number

number

number

number

[0049] The present invention allows a very effective implementation of the redundancy principle, with fewer constraints than traditional lockstep and limited network and computation overhead, since signatures can be generated by very short messages, which is again reduced if the signature calculations are performed by a hardware accelerator.

[0050] The signature generation device described in patent FR 2 989 488 B1 therefore provides an efficient implementation of a signature for execution. n =φ), the signature can be computed by the data concentrator itself.

Claims

1. A computer system (1) installed on a carrier, communicating in a network with a data concentrator (2) and a monitor (M) and implementing at least one service critical to the operational safety of said carrier, said critical service being implemented by different respective computers (C 1 , . . . , C m ) at least two instances (δ 1 ,... , δ m ) is redundant, Each computer (C k ) is an instance (δ k ), and - A series of increasing task activation dates (R n ) and a series of corresponding task latest finish dates (D n ) with an increasing set of task activation dates (R n ) and a series of corresponding task latest finish dates (D n ). the internal state of the computer between two successive activations of the service (s n ) backup; - the corresponding activation date (R n The internal state (s) of the computer for each activation (n) of the service starting after n+1 ) is updated by updating the input data (i n ) and read the output data (o n ) and provide it to the data concentrator (2), firstly, by computing the updated internal state and the computed output data (s n+1 , o n ), and secondly, the dependency between the previous internal state and the read input data (s n , i n ) is represented by a transfer function (f); and - The latest end date (D n ) to the instances (δ k ) signature that characterizes the execution [0010] , by a hash chain depending on a hash function (H) on nb bits, and [0025] A relay server (SR) configured to transmit the k ). The critical service is configured to be executed under time control by using The monitor (M) determines whether the instance (δ 1 ,... , δ m ) said signature [0030] A computer system (1) that detects faults by analyzing a

2. The relay server computes the following relationship by recursion for each instance k within each period n through a hash chain using a cryptographic hash function H: [0045] (In the formula, [0050] represents the signature of the instance k in period n+1; [006] represents the signature of the instance k within the time period n; [0070] represents the input data of the service of the instance k within the time period n; and [0080] represents the output data of the service of the instance k within the period n) The system (1) according to claim 1, configured to calculate the signature by:

3. The monitor (M) determines whether the instance (δ 1 , . . . , δ m Signature of [0090] is the latest end date of the current period (D n 3. The system (1) according to claim 1 or 2, configured to detect a temporal disturbance if the signal is not received before .

4. The monitor (M) determines whether the instance (δ 1 , . . . , δ m ) one signature [0010] But other signatures ##EQU00011## In order to detect an operational failure when the relay server (SR 1 , . . . , SR m The system (1) according to any one of claims 1 to 3, configured to compare the signatures received from the respective authentication servers.

5. The system (1) according to any one of claims 1 to 4, wherein if the number of instances of the service is at least equal to three and less than half of the instances are faulty, the monitor (M) is configured to take a majority vote among the signatures received in time and provide a majority signature indicating an operational instance that is different from a majority signature indicating a faulty instance, and to signal the operational and faulty instances to the rest of the system such that transmission of the faulty instances is interrupted on the data concentrator (2).

6. The instance lasts for period n d If a failure is detected within the period n, the computer hosting the failed instance applies the transfer function to possibly delay the schedule to the corresponding latest end date, and d and resuming the failed instance from said correct memory state, and d is configured to feed back the input data from the nth latest end date (D n 6. The system of claim 5, further configured to report the failed instance as operational again if, before applying the transfer function to inputs up to period n after the period nd, the signatures are again equal.

7. Relay Server (SR k ) corresponds to the corresponding computer (C k The system according to any one of claims 1 to 6, wherein the system is a software server implemented on the

8. Relay Server (SR k 7. The system according to claim 1, wherein the data concentrator (2) is a hardware server implemented in the data concentrator (2).

9. The relay server (SR) has a lower passband and higher reliability than that of the data concentrator (2). k The system of any one of claims 1 to 8, further comprising a network (RI) independent of said data concentrator for transmitting said signature by means of a RI.

10. A method for managing at least one service critical to the operational safety of a computer system according to any one of claims 1 to 9, installed on a carrier, said critical service being managed by each of the different computers (Cs) connected to a network. 1 , . . . , C m ) at least two instances (δ 1 ,... , δ m ) is redundant, The instance (δ) of the critical service k Each implementation of - A series of increasing task activation dates (R n ) and a series of corresponding task latest finish dates (D n ) with an increasing set of task activation dates (R n ) and a series of corresponding task latest finish dates (D n ). the internal state of the computer between two successive activations of the service (s n ) backup; - the corresponding activation date (R n The internal state (s) of the computer for each activation (n) of the service starting after n+1 ) is updated by updating the input data (i n ) and read the output data (o n ) and provide it to the data concentrator (2), firstly, by computing the updated internal state and the computed output data (s n+1 , o n ), and secondly, the dependency between the previous internal state and the read input data (s n , i n ) is represented by a transfer function (f); and - The latest end date (D) from the start (0) of the system to the present by the relay server n ) to the instances (δ k ) signature that characterizes the execution ##EQU00012## Calculation of , by a hash chain depending on a hash function (H) on nb bits and signing said , to a monitor (M). ##EQU00013## Send Using The detection of a failure by the monitor (M) is performed by the instance (δ 1 , . . . , δ m ) said signature ##EQU00014## The method is carried out by analyzing

Citation Information

Patent Citations

  • Method and system for multiprocessing

    JP2002312190A

  • Multi-core microcontroller having comparator for collating processing result

    JP2010113388A

  • Cloud control system provided with plurality of arithmetic servers, scheduling method of control program of the same and redundancy method of the plurality of arithmetic servers

    JP2015219896A

  • Event batching, output sequencing, and log based state storage in continuous query processing

    US20170116089A1

  • Enhancing diagnostic capabilities of computing systems by combining variable patrolling API and comparison mechanism of variables

    US20190235448A1