Data transmission method
By reducing data dimensionality, the method addresses transmission challenges, allowing real-time infraction detection and storage for autonomous vehicles, enhancing training and compliance monitoring.
Patent Information
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- OXA AUTONOMY LTD
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-06
AI Technical Summary
Autonomous vehicles generate large volumes of data that are difficult to transmit over existing communication networks like 4G, making it challenging to monitor and store this data for training and regulatory compliance.
A computer-implemented method that reduces the dimensionality of data descriptors using machine learning models, enabling real-time infraction monitoring and alerting in autonomous vehicles.
Enables real-time infraction detection and response in autonomous vehicles, facilitating data storage for training and compliance verification.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
FIELD
[01] The subject-matter of the present disclosure relates to computer-implemented methods of transmitting data over a communication network, the data representing a scene in which an autonomous vehicle operates, and transitory or non-transitory computer readable media including instructions for performing the method. BACKGROUND
[02] Typically, an autonomous vehicle, AV, captures data relating to scenes in which the AV operates. It would be desirable to transmit this data over a communication network, e.g. a wireless network, so that it can be monitored and stored for various purposes including training autonomy stacks, monitoring regulatory compliance, etc. However, the data relating to scenes are typically very large, and thus not possible to be transmitted over communications networks such as 4G networks.
[03] It is an aim of the present invention to address such problems and improve on the prior art. SUMMARY
[04] According to an aspect of the present disclosure, there is provided a computer-implemented method of monitoring a scene captured by an autonomous vehicle, the computer-implemented method comprising: receiving, by a remote monitoring station, data transmitted from the autonomous vehicle; inputting the data to a machine learning model to predict if the autonomous vehicle experiences an infraction in the scene; storing the data and any predicted infractions in storage at the remote monitoring station; and transmitting an alert to the autonomous vehicle in response to predicting any infractions, wherein the data includes a plurality of descriptors, each descriptor associated with a features in the scene, and a plurality of reduced dimensionality descriptors.
[05] Using the remote monitoring station to predict if an AV encounters an infraction nables infraction monitoring over a fleet of vehicles in real time. The alert transmitted to the AV enables the AV to take action in response to any missed infractions, if necessary. In addition, the AV may be able to store the predicted infraction when no action was taken in response, so the infraction can be logged as a bug of the autonomy stack of the AV. An infraction may be understood to mean an incident involving the AV and other elements of the scene, e.g. a collision, a near collision, etc.
[06] In an embodiment, the features include at least one lane, and wherein the reduced dimensionality descriptors includes a reduced set of coordinate samples for the lane, and wherein the descriptors includes a classification of a lane type.
[07] In an embodiment, the classification of a lane type includes one or an AV lane, an adjacent ongoing lane, an adjacent oncoming lane, a bike lane, a bus lane, an on-ramp lane, and off-ramp lane, a taxi lane, and a tram lane.
[08] In an embodiment, the plurality of features includes at least one object, wherein the descriptor of the at least one object includes: coordinates of the object; a length of the object; a width of the object; a pose of the object; and a velocity of the object.
[09] In an embodiment, the reduced dimensionality descriptor of the at least one object comprises: a numerical string identifying the object; and an object type.
[10] In an embodiment, the object type is a classification of object type selected from a list of object types including: a pedestrian, a cyclist, a vehicle, a motorcyclist, and other.
[11] In an embodiment, the plurality of features includes parameters associated with the AV, wherein a descriptor of the parameters comprises: a length of the AV; a width of the AV; and a speed of the AV.
[12] In an embodiment, the plurality of features includes a planned trajectory of the AV, wherein the reduced dimensionality descriptor for the planned trajectory comprises: coordinates of the planned trajectory; a spline associated with the coordinates, and wherein the descriptor for the planned trajectory comprises: a speed of the AV at each coordinate.
[13] In an embodiment, the plurality of features includes an autonomy state of the AV, where the reduced dimensionality descriptor of the autonomy state includes and enumerated list of autonomy states, and where the autonomy states include: an autonomous mode, and a non-autonomous mode.
[14] In an embodiment, the plurality of features includes a controller error, wherein the descriptor of the controller error includes a difference in actual speed and command speed, and a difference in actual steering angle and commanded steering angle.
[15] In an embodiment, the plurality of features includes a semantic scene element, wherein the descriptor for the semantic scene element comprises: coordinates for the semantic scene element; and a semantic type classification of the semantic scene element, wherein the classification of the semantic scene element is one of: a traffic light, a yield sign, or a stop sign.
[16] In an embodiment, the plurality of features includes junctions, and wherein the descriptor for the junction includes: a lane identifier; coordinates defining a perimeter of the junction associated with the lane; a control type; and a transition type.
[17] In an embodiment, the control type is selected from a list of control types including a traffic light control, a pedestrian control a keep clear sign, a stop sign, a give way sign, and a precedence.
[18] In an embodiment, the transition type is selected from a list of transition types including straight on, left turn, merge, right turn, merge, and a blank.
[19] According to an aspect of the present disclosure, there is provided a computer-implemented method of training a machine learning model to predict an infraction from data transmitted from an autonomous vehicle, the data describing a scene captured by the autonomous vehicle, the computer-implemented method comprising: inputting the data transmitted from the autonomous vehicle to the machine learning model to predict any infractions involving the autonomous vehicle, the data transmitted from the autonomous vehicle having descriptors and reduced dimensionality descriptors; calculating an error between the predictions of the machine learning model and labels which label infractions in the data; and optimising parameters of the machine learning model to minimise the error. BRIEF DESCRIPTION OF DRAWINGS
[20] The subject-matter of the present disclosure is best described with reference to the accompanying figures, in which:
[21] Figure 1 shows a schematic diagram of an AV and a remote monitoring station, according to one or more embodiments;
[22] Figure 2 shows a schematic plan view of the AV traversing a lane, according to one or more embodiments;
[23] Figure 3 shows a schematic plan view with sampled coordinates of the lane from Figure 2, according to one or more embodiments;
[24] Figure 4 shows the schematic plan view similar to Figure 3 using a reduced set of coordinates for dimensionality reduction, according to one or more embodiments;
[25] Figure 5 shows a schematic plan view of the AV traversing a different lane, according to one or more embodiments;
[26] Figure 6 shows a schematic plan view of the AV and sampled coordinates of its planned trajectory, according to one or more embodiments;
[27] Figure 7 shows a schematic plan view similar to Figure 6 using a reduced set of coordinates for dimensionality reduction, according to one or more embodiments;
[28] Figure 8 shows a schematic plan view similar to Figure 7 with a spline fitted to the reduced set of coordinates to reduce information loss, according to one or more embodiments;
[29] Figure 9 shows a schematic plan view of the AV encountering a junction, according to one or more embodiments;
[30] Figure 10 shows a flow chart summarising a computer implemented method of transmitting data over a communication network, according to one or more embodiments;
[31] Figure 11 shows a flow chart summarising scene monitoring using an ML model to predict infractions and transmitting an alert in response, according to one or more embodiments; and
[32] Figure 12 shows a block diagram of training a machine learning model to predict an infraction from data transmitted from an AV, according to one or more embodiments. DESCRIPTION OF EMBODIMENTS
[33] Although the example embodiments have been described with reference to the components, modules and units discussed herein, such functional elements may be combined into fewer elements or separated into additional elements. Various combinations of optional features have been described herein, and it will be appreciated that described features may be combined in any suitable combination. In particular, the features of any one example embodiment may be combined with features of any other embodiment, as appropriate, except where such combinations are mutually exclusive. Throughout this specification, the term “comprising” or “comprises” means including the component(s) specified but not to the exclusion of the presence of others.
[34] The embodiments described herein may be embodied as sets of instructions stored as electronic data in one or more storage media. Specifically, the instructions may be provided on a transitory or non-transitory computer-readable media. When executed by the processor, the processor is configured to perform the various methods described in the following embodiments. In this way, the methods may be computer-implemented methods. In particular, the processor and a storage including the instructions may be incorporated into a vehicle. The vehicle may be an autonomous vehicle (AV).
[35] Whilst the following embodiments provide specific illustrative examples, those illustrative examples should not be taken as limiting, and the scope of protection is defined by the claims. Features from specific embodiments may be used in combination with features from other embodiments without extending the subject-matter beyond the content of the present disclosure.
[36] With reference to Figure 1, an AV 10 may include a plurality of sensors 12. The sensors 12 may be mounted on a roof of the AV 10, or integrated into the bumpers, grill, bodywork, etc. The sensors 12 may be communicatively connected to a computer 14. The computer 14 may be onboard the AV 10. The computer 14 may include a processor 16 and storage 18. The memory may include the non-transitory computer-readable media described above. Alternatively, the non-transitory computer-readable media may be located remotely and may be communicatively linked to the computer 14 via the cloud 20. The computer 14 may be communicatively linked to one or more actuators 22 for control thereof to move the AV 10. The actuators may include, for example, a motor, a braking system, a power steering system, etc.
[37] The computer 14 includes an autonomy stack for controlling the AV 10 stored the storage 18. The autonomy stack 34 may control the AV 10 in response to the sensor data. To achieve this, the autonomy stack may include one or more machine learning models. The one or more machine learning models may include an end-to-end model that is trained to provide control commands to actuators of the AV 10 in response to the sensor data. The one or more machine learning models may include machine learning models respectively responsible for perception, planning, and control. This may be in addition to the end-to-end model or as an alternative to the end-to-end model. Perception functions may include object detection and classification based on sensor data. Planning functions may include object tracking and trajectory generation. Control functions including setting control instructions for one or more actuators 22 of the AV 10 to move the AV 10 according to the trajectory.
[38] The sensors 12 may include various sensor types. Examples of sensor types include LiDAR sensors, RADAR sensors, and cameras. Each sensor type may be referred to as a sensor modality. Each sensor type may record data associated with the sensor modality. For example, the LiDAR sensor may record LiDAR modality data.
[39] The data may capture various scenes that the AV 10 encounters. For example, a scene may be a visible scene around the AV 10 and may include roads, buildings, weather, objects (e.g. other vehicles, pedestrians, animals, etc.), etc.
[40] The communication via the cloud 20 may be over a communication network. The communication network may be a telecommunication network, e.g. 4G, 5G, or similar. The AV 10 may communicate with a remote monitoring station 24. The remote monitoring station 24 includes a computer comprising at least one processor 26 and storage 28. The storage may have a machine learning model stored thereon for assessing whether data associated with a scene, and received from the AV 10 over the communications network, describes any infractions between the AV 10 and another actor in the scene.
[41] The computer of the AV 10 may be configured to perform a computer-implemented method of transmitting data over the communication network. The data represents a scene in which the AV 10 operates, as described above. The AV 10 may thus be considered an ego-vehicle.
[42] The computer-implemented method may be called simply, a method.
[43] The method comprises capturing, from at least one of the sensors 12, the scene. For example, the scene may be captured using images and features in the images identified using a machine learning algorithm such as a you only look once, YOLO, algorithm. The scene may also be captured using LiDAR sensor, to form a point cloud. Similarly, a point cloud can be created using RADAR points from the radar sensor.
[44] The method also includes constructing a descriptor for each feature, of a plurality of features, in the scene. The descriptor may be an alphanumerical string, for example.
[45] The method also includes storing the descriptors in the storage 18 of the AV, reducing a dimensionality of at least one of the descriptors and transmitting data over the communication network, the data including the reduced dimensionality descriptors. The data transmitted over the communications network may also include any descriptors that have not been reduced dimensionally.
[46] With reference to Figure 2, the plurality of features includes at least one lane 30.
[47] With reference to Figure 3, constructing the descriptor for the at least one lane comprises sampling coordinates 32 for the lane 30, sampling widths of the lane 30 along the lane, and classifying the at least one lane as a lane type. The coordinates of the lane 30 may be sampled along opposing edges of the lane. The coordinates may be cartesian coordinates made up of an x ordinate and a y ordinate. The coordinates may be relative to an origin located at the AV 10. In other words, the coordinates may be relative to the AV 10. Alternatively, the coordinates may be relative to an origin at a point on a map including routes along which the AV 10 travels.
[48] With reference to Figure 4, reducing the dimensionality of at least one of the descriptors comprises reducing a number of sampled coordinates 32 for the lane 30. For example, if the lane has 5 coordinates around a bend, those may be reduced to the first, third, and fifth coordinates, thus reducing the number of coordinates to 3 around the bend. Along a straight lane, the start and end point of the lane may be retained, and any intermediate points may be deleted. The specific example in Figure 3 shows 11 coordinates at each edge of the lane 32. The reduced dimensionality lane 30 in Figure 4 is represented using a reduced set of coordinates, e.g. 4 at each edge of the lane.
[49] The lane type may be classified as one of an AV lane, an adjacent ongoing lane, an adjacent oncoming lane, a bike lane, a bus lane, an on-ramp lane, an off-ramp lane, a taxi lane, or a tram lane.
[50] With reference to Figure 5, the plurality of features includes at least one object 34. Constructing the descriptor for the at least one object 34 comprises capturing coordinates of the object 34, capturing a length of the object 34, capturing a width of the object 34, capturing a pose of the object 34, capturing a velocity of the object 34, generating a numerical string to identify the object 34, and classifying the object as an object type.
[51] The numerical string may a number string having a first integer, a decimal point, and a plurality of further integers after the decimal point. It is possible to reduce the dimensionality of the at least one descriptor. This can be achieves by reducing a length of the numerical string by selecting a first mentioned predetermined number of characters in the numerical string. For example, the predetermined number may be three, four, five, or another number. Where the first four numbers are to be selected, and the numerical string identifying a cyclist is 1.43245693984, the reduced dimensionality version may be 1.432.
[52] The dimensionality may also be reduced by reducing a name size of the object. For example, the name of the object may be CLASS_TYPE_PEDESTRIAN, which may be reduced to PEDESTRIAN by retaining only the unique identifier portion. The object type may be classified as an object selected from a list of objects including a pedestrian, a cyclist, a vehicle, a motorcyclist, and other types of objects. The objects thus may be other actors that are able to move in the scene.
[53] The plurality of features may also include parameters associated with the AV 10. Constructing the descriptor for the parameters associated with the AV 10 may comprise capturing a length of the AV, capturing a width of the AV, and capturing a speed of the AV.
[54] With reference to Figure 6, the plurality of features includes a planned trajectory 36 of the AV 10. Constructing a descriptor for the planned trajectory of the AV 10 comprises sampling coordinates 36 of the planned trajectory. The descriptor may also include speed samples of the AV at each sampled coordinate.
[55] With reference to Figure 7, in a similar way to the path coordinates, the dimensionality of the planned trajectory may be reduced by reducing a number of coordinates. For example, one in three coordinates may be retained and the other two may be discarded, when the trajectory is around a bend. When the trajectory is a straight line, the start and end points of that section of the trajectory only may be retained.
[56] With reference to Figure 8, a spline 38 may be fitted to the reduced set of the coordinates 36. The spline can be included in the reduced dimensionality descriptor to reduce information loss. The spline is particularly useful for the planned trajectories that turn around a bend, for example.
[57] The plurality of descriptors may also include an autonomy state of the AV. Constructing the descriptor for each feature in the scene comprises generating a list of autonomy states from a list including: autonomous mode, and non-autonomous mode. This may be reduced in dimensionality by enumerating (ENIIM) the list of autonomy states.
[58] The plurality of features includes a controller error. Constructing a descriptor for the controller error comprises calculating a difference in actual speed and commanded speed, and calculating a difference in actual steering angle and commanded steering angle.
[59] With further reference to Figure 5, the plurality of features includes a semantic scene element 42. The descriptor for the semantic scene element may constructed by capturing the coordinates for the semantic scene element and classifying the semantic scene element as a semantic type. The semantic type may refer to non-actor elements in the scene. The non-actor types may be features that do not change position, so they may be static rather than dynamic like objects. The semantic scene element may be classified as one of a traffic light, a yield sign, and a stop sign.
[60] The descriptor may record the semantic type for example as SEMANTIC_TYPE_TRAFFIC LIGHT, SEMANTIC_TYPE_YIELD SIGN, etc. The reduced dimensionality descriptor may reduce a size of the semantic type label. For example, the label may be reduced to TRAFFIC LIGHT, YIELD SIGN, etc.
[61] With reference to Figure 9, the plurality of features may include junctions. The descriptor may be constructed by creating a lane identifier, sampling coordinates 44 defining a perimeter 46 of the junction, generating a control type, and generating a transition type.
[62] The lane identifier may be a numerical string, similar to the numerical string described above.
[63] The control type may refer to a semantic scene element 42 used to control the junction. For example, the control type is selected from a list of control types including traffic light control, pedestrian control, e.g. a pelican crossing or zebra crossing, a keep clear sign, a stop sign, a give way sign, and a precedence. The precedence may be derive from road laws and rules, e.g. at unmarked cross-roads.
[64] The transition type may refer to how the AV 10 should transition through the junction. The transition types may include straight on, left turn, merge, right turn, merge, and a blank.
[65] At the remote monitoring station 24, the scene may be monitored. This may be done using a computer-implemented method, performed by the processor and storage of the remote monitoring station 24, of monitoring a scene captured by the AV 10.
[66] The method comprises reciting, by the remote monitoring station 24, data transmitted from the AV. The data is the data described above including descriptors and reduced dimensionality descriptors of the various elements describing the scene. The data may be transmitted periodically, e.g. at a frequency of 1 data packet per second, or more, e.g. 20 data packets per second, or fewer, e.g. one data packet per minute. The term data packet may be used to mean a complete set of descriptors and reduced dimensionality descriptors that describe a snapshot of a scene captured by the AV.
[67] The method also comprises inputting the data to a machine learning model to predict if the autonomous vehicle experiences an infraction in the scene. An infraction may be an even involving the AV 10 and another element of the scene, e.g. an object or semantic scene element, which causes the AV 10 to take evasive action, e.g. perform a minimal risk manoeuvre MRM such as an emergency stop or pulling to the side of a road. Evasive action is not limited to such extreme scenarios and may also mean deviating from a planned trajectory.
[68] The method also includes storing the data and any predicted infractions in the storage 28 of the remote monitoring station. In this way, the data can be retrieved subsequently for training future iterations of the autonomy stack, for regulatory compliance checks, and for explainability of the autonomy stack, etc.
[69] The method also comprises transmitting an alert to the autonomous vehicle in response to predicting any infractions. The alert is transmitted back over the same communications network. The alert will alert the AV 10 to the fact that an infraction has been detected. If the AV 10 actually detected the same infraction, the autonomy stack may be validated as working correctly. However, if the AV 10 did not detect an infraction, and so took no evasive action, the autonomy stack can register this discrepancy so that the autonomy stack can be tested and adapted where necessary.
[70] The machine learning model may be trained to predict the infraction from the data transmitted from the AV 10. This may be achieved using a computer-implemented method which may be performed at the remote monitoring station 24, or offline at another computer and then the machine learning model saved to the storage 28 of the remote monitoring station 24.
[71] The method comprises inputting data from the AV to the machine learning model to predict any infractions involving the autonomous vehicle. Previously, labels have been added to the data to indicate infractions. The labels may be generated manually via human input or automatically using another machine learning model. The method comprises calculating an error between the predictions of the machine learning model and labels which label infractions in the data. The method also comprises optimising parameters of the machine learning model to minimise the error.
[72] The machine learning model may be a neural network. The parameters of the neural network may be its weights. The optimising process may use algorithms such as back-propagation and gradient descent.
[73] With reference to Figure 10, the computer-implemented method of transmitting data over a communication network, the data representing a scene in which an autonomous vehicle, AV, operates, is summarised as including: capturing 100, from at least one sensor of the AV, a scene; constructing 102 a descriptor for each feature, of a plurality of features, in the scene; storing 104 the descriptors in storage of the AV; reducing 106 a dimensionality of at least one of the descriptors; and transmitting 108 data over the communication network, the data including the reduced dimensionality descriptors.
[74] With reference to Figure 11, the computer-implemented method of monitoring a scene captured by the AV 10 may be summarised as comprising: receiving 200, by a remote monitoring station, data transmitted from the autonomous vehicle; inputting 202 the data to a machine learning model to predict if the autonomous vehicle experiences an infraction in the scene; storing 204 the data and any predicted infractions in storage at the remote monitoring station; and 206 transmitting an alert to the autonomous vehicle in response to predicting any infractions, wherein the data includes a plurality of descriptors, each descriptor associated with a features in the scene, and a plurality of reduced dimensionality descriptors.
[75] With reference to Figure 12, the computer-implemented method of training a machine learning model to predict an infraction from data transmitted from an autonomous vehicle may be summarised as comprising: inputting 300 the data transmitted from the autonomous vehicle to the machine learning model to predict any infractions involving the autonomous vehicle, the data transmitted from the autonomous vehicle having descriptors and reduced dimensionality descriptors; calculating 302 an error between the predictions of the machine learning model and labels which label infractions in the data; and optimising 304 parameters of the machine learning model to minimise the error.
[76] The following clauses may be useful for understanding details related to the present disclosure.
[77] Clause 1. A computer-implemented method of transmitting data over a communication network, the data representing a scene in which an autonomous vehicle, AV, operates, the computer-implemented method comprising: capturing, from at least one sensor of the AV, a scene; constructing a descriptor for each feature, of a plurality of features, in the scene; storing the descriptors in storage of the AV; reducing a dimensionality of at least one of the descriptors; and transmitting data over the communication network, the data including the reduced dimensionality descriptors.
[78] Clause 2. The computer-implemented method of Clause 1, wherein the plurality of features includes at least one lane, wherein constructing the descriptor for the at least one lane comprises: sampling coordinates for the lane; sampling widths of the lane along the lane; and classifying the at least one lane as a lane type.
[79] Clause 3. The computer-implemented method of Clause 2, wherein reducing the dimensionality of at least one of the descriptors comprises: reducing a number of sampled coordinates for the lane.
[80] Clause 4. The computer-implemented method of Clause 2 or Clause 3, wherein classifying the at least one lane as a lane type comprises: classifying the at least one lane as an AV lane, an adjacent ongoing lane, an adjacent oncoming lane, a bike lane, a bus lane, an on-ramp lane, and off-ramp lane, a taxi lane, and a tram lane.
[81] Clause 5. The computer-implemented method of any preceding clause, wherein the plurality of features includes at least one object, wherein constructing the descriptor for the at least one object comprises: capturing coordinates of the object; capturing a length of the object; capturing a width of the object; capturing a pose of the object; capturing a velocity of the object; generating a numerical string to identify the object; and classifying the object as an object type.
[82] Clause 6. The computer-implemented method of Clause 5, wherein reducing the dimensionality of the at least one descriptor comprises: reducing a length of the numerical string by selecting a first mentioned predetermined number of characters in the numerical string; and reducing a name size of the object type.
[83] Clause 7. The computer-implemented method of Clause 5 or Clause 6, wherein classifying the object as an object type comprises: classifying the object as an object selected from a list of objects including a pedestrian, a cyclist, a vehicle, a motorcyclist, and other.
[84] Clause 8. The computer-implemented method of any preceding clause, wherein the plurality of features includes parameters associated with the AV, wherein constructing the descriptor for the parameters associated with the AV comprises: capturing a length of the AV; capturing a width of the AV; and capturing a speed of the AV.
[85] Clause 9. The computer-implemented method of any preceding clause, wherein the plurality of features includes a planned trajectory of the AV, wherein constructing a descriptor for the trajectory of the AV comprises: sampling coordinates of the planned trajectory; and sampling a speed of the AV at each sampled coordinate.
[86] Clause 10. The computer-implemented method of Clause 9, wherein reducing the dimensionality of the at least one descriptor comprises: reducing a number of samples of the coordinates for the planned trajectory.
[87] Clause 11. The computer-implemented method of Clause 10, further comprising computing a spline for the planned trajectory using the reduced number of samples.
[88] Clause 12. The computer-implemented method of any preceding clause, wherein the plurality of features includes an autonomy state of the AV, wherein constructing a descriptor for each feature, of a plurality of features, in the scene comprises: generating a list of autonomy states from a list including: autonomous mode, and non-autonomous mode.
[89] Clause 13. The computer-implemented method of Clause 11, wherein reducing the dimensionality of the list of autonomy states comprises: enumerating the list of autonomy states.
[90] Clause 14. The computer-implemented method of any preceding clause, wherein the plurality of features includes a controller error, wherein constructing a descriptor for the controller error comprises calculating a difference in actual speed and commanded speed, and calculating a difference in actual steering angle and commanded steering angle.
[91] Clause 15. The computer-implemented method of any preceding clause, wherein the plurality of features includes a semantic scene element, wherein constructing a descriptor for the semantic scene element comprises: capturing coordinates for the semantic scene element; and classifying the semantic scene element as a semantic type.
[92] Clause 16. The computer-implemented method of Clause 14, wherein classifying the semantic scene element as a semantic type comprises: classifying the semantic scene element as one of: a traffic light, a yield sign, and a stop sign.
[93] Clause 17. The computer-implemented method of Clause 14 or Clause 15, wherein reducing the dimensionality of the semantic scene element comprises: reducing a length of the semantic type.
[94] Clause 18. The computer-implemented method of any preceding clause, wherein the plurality of features includes junctions, and wherein constructing a descriptor comprises constructing a descriptor for a junction by: creating a lane identifier; sampling coordinates defining a perimeter of a junction; generating a control type; and generating a transition type.
[95] Clause 19. The computer-implemented method of Clause 18, wherein the control type is selected from a list of control types including a traffic light control, a pedestrian control a keep clear sign, a stop sign, a give way sign, and a precedence.
[96] Clause 20. The computer-implemented method of Clause 18 or Clause 19, wherein the transition type is selected from a list of transition types including straight on, left turn, merge, right turn, merge, and a blank.
[97] Clause 21. The computer-implemented method of any preceding clause, wherein the data transmitted over the communications network also includes any descriptors that have not been reduced dimensionally.
[98] Clause 22. A transitory, or non-transitory, computer-readable medium, having instructions stored thereon that when executed by at least one processor, causes the at least one processor to perform the computer-implemented method of any preceding clause.
[99] While the invention has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive; the invention is not limited to the disclosed embodiments.
[100] Other variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. A single processor or other unit may fulfil the functions of several items recited in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measured cannot be used to advantage. Any reference signs in the claims should not be construed as limiting the scope.
Claims
1. A computer-implemented method of monitoring a scene captured by an autonomous vehicle, the computer-implemented method comprising: receiving, by a remote monitoring station, data transmitted from the autonomous vehicle; inputting the data to a machine learning model to predict if the autonomous vehicle experiences an infraction in the scene; storing the data and any predicted infractions in storage at the remote monitoring station; and transmitting an alert to the autonomous vehicle in response to predicting any infractions, wherein the data includes a plurality of descriptors, each descriptor associated with a features in the scene, and a plurality of reduced dimensionality descriptors.
2. The computer-implemented method of Claim 1, wherein the features includes at least one lane, and wherein the reduced dimensionality descriptors includes a reduced set of coordinate samples for the lane, and wherein the descriptors includes a classification of a lane type.
3. The computer-implemented method of Claim 2, wherein the classification of a lane type includes one or an AV lane, an adjacent ongoing lane, an adjacent oncoming lane, a bike lane, a bus lane, an on-ramp lane, and off-ramp lane, a taxi lane, and a tram lane.
4. The computer-implemented method of any preceding claim, wherein the plurality of features includes at least one object, wherein the descriptor of the at least one object includes: coordinates of the object; a length of the object; a width of the object; a pose of the object; and a velocity of the object.
5. The computer-implemented method of Claim 4, wherein the reduced dimensionality descriptor of the at least one object comprises: a numerical string identifying the object; and an object type.
6. The computer-implemented method of Claim 5, wherein the object type is a classification of object type selected from a list of object types including: a pedestrian, a cyclist, a vehicle, a motorcyclist, and other.
7. The computer-implemented method of any preceding claim, wherein the plurality of features includes parameters associated with the AV, wherein a descriptor of the parameters comprises: a length of the AV; a width of the AV; and a speed of the AV.
8. The computer-implemented method of any preceding claim, wherein the plurality of features includes a planned trajectory of the AV, wherein the reduced dimensionality descriptor for the planned trajectory comprises: coordinates of the planned trajectory; a spline associated with the coordinates, and wherein the descriptor for the planned trajectory comprises: a speed of the AV at each coordinate.
9. The computer-implemented method of any preceding claim, wherein the plurality of features includes an autonomy state of the AV, where the reduced dimensionality descriptor of the autonomy state includes and enumerated list of autonomy states, and where the autonomy states include: an autonomous mode, and a non-autonomous mode.
10. The computer-implemented method of any preceding claim, wherein the plurality of features includes a controller error, wherein the descriptor of the controller error includes a difference in actual speed and command speed, and a difference in actual steering angle and commanded steering angle.
11. The computer-implemented method of any preceding claim, wherein the plurality of features includes a semantic scene element, wherein the descriptor for the semantic scene element comprises: coordinates for the semantic scene element; and a semantic type classification of the semantic scene element, wherein the classification of the semantic scene element is one of: a traffic light, a yield sign, or a stop sign.
12. The computer-implemented method of any preceding claim, wherein the plurality of features includes junctions, and wherein the descriptor for the junction includes: a lane identifier; coordinates defining a perimeter of the junction associated with the lane; a control type; and a transition type.
13. The computer-implemented method of Claim 12, wherein the control type is selected from a list of control types including a traffic light control, a pedestrian control a keep clear sign, a stop sign, a give way sign, and a precedence.
14. The computer-implemented method of Claim 12 or Claim 13, wherein the transition type is selected from a list of transition types including straight on, left turn, merge, right turn, merge, and a blank.
515. A computer-implemented method of training a machine learning model to predict an infraction from data transmitted from an autonomous vehicle, the data describing a scene captured by the autonomous vehicle, the computer-implemented method comprising: inputting the data transmitted from the autonomous vehicle to the 10 machine learning model to predict any infractions involving the autonomous vehicle, the data transmitted from the autonomous vehicle having descriptors and reduced dimensionality descriptors; calculating an error between the predictions of the machine learning model and labels which label infractions in the data; and optimising parameters of the machine learning model to minimise the error.1519
Citation Information
Patent Citations
Picture transmission method and device, equipment and storage medium
CN113592003A