Emergency instruction automatic generation system and method for road emergency handling, terminal and medium

By adopting multimodal data processing and neural network models in the handling of road emergencies, the problems of unreasonable evaluation delay and emergency command issuance in the existing technology are solved, and more efficient and accurate emergency response is achieved.

CN120108172APending Publication Date: 2025-06-06SHANGHAI SANSI ELECTRONICS ENG +4
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311662443.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-05
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the handling of road emergencies, the problem of delay in assessment, insufficient accuracy, and unreasonable and incorrect issuance of emergency instructions.

Method used

The multimodal data acquisition and processing module is adopted to process and judge multimodal data through neural network model and hierarchical structure model to generate emergency instructions.

Benefits of technology

It improves the timeliness and accuracy of road emergencies, ensures the rationality and correctness of emergency instructions, and reduces personnel and property losses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108172A_ABST
    Figure CN120108172A_ABST
Patent Text Reader

Abstract

According to the emergency instruction automatic generation system and method for road emergency disposal, the terminal and the medium provided by the invention, different characteristics of road sudden dangerous events can be comprehensively captured by utilizing multi-modal data such as text data, audio data, image data and video data through a generative artificial intelligence technology. Performing data representation mapping on the multi-modal data to complete classification and risk level evaluation of the law emergencies; meanwhile, qualitative judgment of sudden road dangerous events is achieved in combination with technologies such as a hierarchical multi-label classification network, and the timeliness and accuracy of judgment can be effectively improved; through a generative artificial intelligence technology, a feedforward neural network, a multi-head self-attention mechanism, an encoder, a decoder structure network and other technologies, a corresponding emergency instruction is generated according to an evaluation result of a road sudden dangerous event, and the reasonability and correctness of emergency instruction issuing can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of road emergency event identification, and in particular to an emergency command automatic generation system, method, terminal and medium for handling road emergency events. Background Art

[0002] At present, in response to sudden dangerous incidents on the road such as sudden traffic accidents and natural disasters, we mainly rely on law enforcement personnel and management personnel to rush to the scene of the incident, conduct dispatch and command, and complete emergency response.

[0003] However, due to the randomness of the time of occurrence, the uncertainty of the location and the unknown degree of harm of sudden dangerous events on the road, this method of manual emergency response has a lot of uncertainties in terms of timeliness and accuracy, which will have a great negative impact on the degree of casualties and property losses, the scope of harm and the incidence of secondary disasters.

[0004] Especially in high-pressure and emergency situations, law enforcement officers or managers with different experiences will make different decisions and judgments and issue different emergency instructions; and when law enforcement officers or managers lack timely communication, they may issue contradictory emergency instructions at the same time. Summary of the invention

[0005] In view of the above-mentioned shortcomings of the prior art, the present invention provides a system, method, terminal and medium for automatically generating emergency instructions for handling road emergencies, which are used to solve the problems of delayed road emergency assessment, insufficient accuracy, and unreasonable and incorrect issuance of emergency instructions in the prior art.

[0006] To achieve the above-mentioned purpose and other related purposes, the first aspect of the present application provides an automatic generation system of emergency instructions for handling road emergencies, including: a multimodal data acquisition and processing module, which is used to acquire multimodal data of road emergencies in real time and pre-process the multimodal data; a road emergency assessment module, which is used to process and judge the pre-processed multimodal data using a neural network model and a hierarchical structure model to obtain an assessment result of the road emergency; and an emergency instruction generation module, which is used to generate corresponding emergency instructions based on the assessment result of the road emergency.

[0007] In some embodiments of the first aspect of the present application, the system further includes: an emergency instruction upload module; the emergency instruction upload module is used to upload the generated emergency instructions to the cloud to remind the handling of road emergencies.

[0008] To achieve the above-mentioned purpose and other related purposes, the second aspect of the present application provides a method for automatically generating emergency instructions for handling road emergencies, which is applied to the automatic generation system for emergency instructions for handling road emergencies as described above. The method includes: real-time collection of multimodal data of road emergencies, and preprocessing of the multimodal data; processing and judging the preprocessed multimodal data using a neural network model and a hierarchical structure model to obtain an evaluation result of the road emergency; and generating corresponding emergency instructions based on the evaluation result of the road emergency to complete the handling of the road emergency.

[0009] In some embodiments of the second aspect of the present application, the specific process of using a neural network model and a hierarchical model to process and judge the preprocessed multimodal data to obtain the assessment result of the road emergency includes: using a neural network model to perform data representation mapping on the preprocessed multimodal data to obtain the weighted probability distribution of the category and risk level of the road emergency; and using a hierarchical model to judge based on the weighted probability distribution of the category and risk level of the road emergency to obtain the assessment result of the road emergency.

[0010] In some embodiments of the second aspect of the present application, the neural network model includes a linear layer, a nonlinear layer, a classification layer and a fully connected layer, wherein: the linear layer performs a linear transformation on the preprocessed multimodal data to obtain a linear transformation result; the nonlinear layer performs a nonlinear transformation on the linear transformation result to obtain a high-dimensional dense vector; the classification layer uses an activation function to perform classification calculations on the high-dimensional dense vector to obtain the probability distribution of the category and risk level of the road emergency; the fully connected layer performs a weighted operation on the probability distribution of the category and risk level of the road emergency to obtain a weighted probability distribution of the category and risk level of the road emergency.

[0011] In some embodiments of the second aspect of the present application, the specific process of generating corresponding emergency instructions based on the evaluation results of the road emergency includes: using an encoder to calculate and obtain a feature context vector based on the evaluation results of the road emergency; processing the feature context vector in a decoder to obtain a feature weighted context vector; calculating the feature weighted context vector to generate a probability distribution of emergency instructions; and using a decision strategy to select based on the probability distribution of the emergency instructions to generate the optimal emergency instructions.

[0012] In some embodiments of the second aspect of the present application, the encoder includes an embedding layer, a position encoding layer, a first multi-head self-attention layer and a first feedforward neural network layer; wherein: the embedding layer maps the evaluation results of the road emergency to obtain a continuous word embedding vector; the position encoding layer adds position information to the word embedding vector to obtain a sequence representation result of the evaluation result; the first multi-head self-attention layer calculates the sequence representation result to generate a context vector; the first feedforward neural network layer performs nonlinear transformation and feature extraction on the context vector to obtain a feature context vector.

[0013] In some embodiments of the second aspect of the present application, the decoder includes a second multi-head self-attention layer and a second feedforward neural network layer; wherein: the second multi-head self-attention layer calculates the feature context vector to obtain a weighted context vector; the second feedforward neural network layer performs nonlinear transformation and feature extraction on the weighted context vector to obtain a feature weighted context vector.

[0014] To achieve the above-mentioned purpose and other related purposes, the third aspect of the present application provides an electronic terminal, including: a processor and a memory; the memory is used to store computer programs; the processor is used to execute the computer programs stored in the memory, so that the electronic terminal executes the method for automatically generating emergency instructions for handling road emergencies.

[0015] To achieve the above-mentioned purpose and other related purposes, the fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for automatically generating emergency instructions for handling road emergencies.

[0016] As described above, the system, method, terminal and medium for automatically generating emergency instructions for handling road emergencies of the present application have the following beneficial effects:

[0017] (1) This application uses generative artificial intelligence technology and multimodal data, such as text data, audio data, image data, and video data, to comprehensively capture the different characteristics of sudden road hazards. Multimodal data is mapped to complete the classification of sudden road hazards and risk level assessment; at the same time, the qualitative judgment of sudden road hazards is achieved by combining technologies such as hierarchical multi-label classification networks, which can effectively improve the timeliness and accuracy of judgment.

[0018] (2) Through generative artificial intelligence technology, using feedforward neural networks, multi-head self-attention mechanisms, encoder-decoder structure networks and other technologies, the corresponding emergency instructions are generated based on the assessment results of sudden dangerous events on the road, which can effectively improve the rationality and correctness of the issuance of emergency instructions. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 Shown is a structural schematic diagram of an emergency instruction automatic generation system for handling road emergencies in one embodiment of the present application.

[0020] Figure 2 Shown is a flow chart of a method for automatically generating emergency instructions for handling road emergencies in one embodiment of the present application.

[0021] Figure 3 Shown is a flowchart of text data preprocessing in one embodiment of the present application.

[0022] Figure 4 Shown is a schematic diagram of a flow chart of calculating Mel-cepstral coefficients from audio data in one embodiment of the present application.

[0023] Figure 5 Shown is a schematic diagram of the process of performing data enhancement operation on image data in one embodiment of the present application.

[0024] Figure 6 Shown is a schematic diagram of the process of road emergency assessment in one embodiment of the present application.

[0025] Figure 7 Shown is a schematic diagram of a process for performing data representation mapping on preprocessed multimodal data in one embodiment of the present application.

[0026] Figure 8 Shown is a schematic diagram of the linear transformation process in one embodiment of the present application.

[0027] Fig. 9 Shown is a schematic diagram of the process of nonlinear transformation in one embodiment of the present application.

[0028] Fig.10 Shown is a schematic diagram of the result of a high-dimensional dense vector in one embodiment of the present application.

[0029] Fig.11 Shown is a flowchart of multi-classification of the Softmax activation function in one embodiment of the present application.

[0030] Fig.12 Shown is a schematic diagram of the structure of a fully connected layer in one embodiment of the present application.

[0031] Fig.13 Shown is a schematic diagram of a process for generating an emergency instruction in an embodiment of the present application.

[0032] Fig.14 Shown is a flow chart of an encoder in one embodiment of the present application.

[0033] Fig.15Shown is a schematic diagram of the result of relative position encoding in one embodiment of the present application.

[0034] Fig.16 Shown is a flow chart of the multi-head self-attention mechanism in one embodiment of the present application.

[0035] Fig.17 Shown is a schematic diagram of the structure of a feedforward neural network in one embodiment of the present application.

[0036] Fig.18 Shown is a flowchart of a decoder in an embodiment of the present application.

[0037] Fig.19 Shown is a schematic diagram of the structure of an electronic terminal in one embodiment of the present application. DETAILED DESCRIPTION

[0038] The following describes the embodiments of the present application through specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0039] It should be noted that in the following description, with reference to the accompanying drawings, several embodiments of the present application are described in the accompanying drawings. It should be understood that other embodiments may also be used, and mechanical composition, structure, electrical and operational changes may be made without departing from the spirit and scope of the present application. The following detailed description should not be considered restrictive, and the scope of the embodiments of the present application is limited only by the claims of the published patents. The terms used here are only for describing specific embodiments and are not intended to limit the present application. Spatially related terms, such as "upper", "lower", "left", "right", "below", "below", "lower", "above", "upper", etc., may be used in the text to facilitate the description of the relationship between an element or feature shown in the figure and another element or feature.

[0040] In this application, unless otherwise clearly specified and limited, the terms "install", "connect", "connect", "fix", "hold" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0041] Furthermore, as used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless there is an indication to the contrary in the context. It should be further understood that the terms "comprise", "include" indicate the presence of the described features, operations, elements, components, items, kinds, and / or groups, but do not exclude the presence, occurrence or addition of one or more other features, operations, elements, components, items, kinds, and / or groups. The terms "or" and "and / or" used herein are interpreted as inclusive, or mean any one or any combination. Therefore, "A, B or C" or "A, B and / or C" means "any of the following: A; B; C; A and B; A and C; B and C; A, B and C". Exceptions to this definition will only occur when the combination of elements, functions or operations is inherently mutually exclusive in some way.

[0042] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the technical solution in the embodiments of the present invention is further described in detail through the following embodiments and in combination with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the invention.

[0043] Before further describing the present invention in detail, the nouns and terms involved in the embodiments of the present invention are explained. The nouns and terms involved in the embodiments of the present invention are applicable to the following interpretations:

[0044] <1> Intelligent Road Side Unit (RSU): RSU is a key device for the vehicle-road cooperative system to realize the networking and intelligentization of road infrastructure. It undertakes the important task of communication between roads, vehicles and platforms, and is an important communication device for realizing smart transportation. RSU can interconnect data with roadside sensing equipment such as radars and cameras, gather road environment perception information, and realize real-time monitoring of traffic conditions on the platform side. At the same time, RSU can provide vehicles with real-time dynamic traffic information broadcasts, including traffic event information, high-precision positioning, map distribution and updates, etc., which helps reduce driving safety accidents and improve road traffic efficiency.

[0045] <2> Word Embedding is a technique in the field of natural language processing (NLP) and machine learning that is used to map words or tokens (such as words, phrases, symbols) in a text into a continuous vector space. The goal of this technique is to convert discrete text information into continuous real number vectors so that computers can better understand and process text data. These vectors are designed to capture the similarities and associations between words, thereby providing information about the meaning of the words. Word embedding is usually implemented through an embedding matrix, where the number of rows of this matrix is ​​equal to the number of words in the vocabulary, each row corresponds to a word, and the number of columns is usually a pre-defined embedding dimension.

[0046] <3> Natural language processing (NLP) technology mainly studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing is a typical marginal interdisciplinary subject that integrates linguistics, computer science, and mathematics. It is a discipline that uses computer technology to analyze, understand, and process natural language, using computers as a powerful tool for language research, conducting quantitative research on language information with the support of computers, and providing language descriptions that can be used by humans and computers.

[0047] <4> Mel-scale Frequency Cepstral Coefficients (MFCC): MFCC is a commonly used speech feature that can be used to represent the spectral characteristics of speech signals. MFCC is a cepstral parameter extracted in the Mel-scale frequency domain. It is derived from the cepstral spectrum of an audio clip. The difference between the cepstral and the Mel-frequency cepstral is that the frequency band division of the Mel-frequency cepstral is equidistant on the Mel scale, which better approximates the human auditory system than the linearly spaced frequency bands used in the normal logarithmic cepstral. The advantage of MFCC is that it has good human perception properties, can effectively extract key features of speech signals, and is widely used in speech recognition, speech synthesis and other fields.

[0048] <5> Fast Fourier Transformation (FFT): It is obtained by improving the discrete Fourier transform algorithm based on the odd, even, imaginary, and real characteristics of the discrete Fourier transform. The use of this algorithm can greatly reduce the number of multiplications required for the computer to calculate the discrete Fourier transform, especially the more the number of sample points N to be transformed, the more significant the savings in the FFT algorithm calculation.

[0049] <6> Discrete Cosine Transform (DCT): DCT is a mathematical operation closely related to Fourier transform. In the Fourier series expansion, if the expanded function is a real even function, then its Fourier series only contains cosine terms. Discretizing it can derive the cosine transform, so it is called discrete cosine transform.

[0050] <7> Positional encoding is a technique used in natural language processing to process sequence information in text data. The main purpose of this technique is to provide the model with information about the position of elements in the text to help the model better understand the structure and semantics of the text. In NLP tasks such as text classification, machine translation, or text generation, the order of the sequence is often crucial, so positional encoding is very important to consider the positional relationship in the text.

[0051] like Figure 1 , a schematic diagram of the structure of an automatic generation system of emergency instructions for handling road emergencies in an embodiment of the present application is shown. The system includes a multimodal data acquisition and processing module 110, a road emergency assessment module 120 and an emergency instruction generation module 130.

[0052] Among them, the multimodal data acquisition and processing module 110 is used to collect multimodal data of road emergencies in real time and pre-process the multimodal data; the road emergency assessment module 120 is used to process and judge the pre-processed multimodal data using a neural network model and a hierarchical model to obtain the assessment results of the road emergencies; the emergency instruction generation module 130 is used to generate corresponding emergency instructions based on the assessment results of the road emergencies.

[0053] The road emergency assessment module 120 includes a data representation mapping unit 121 and an assessment result generation unit 122; wherein: the data representation mapping unit 121 is used to perform data representation mapping on the preprocessed multimodal data to obtain the weighted probability distribution of the category and risk level of the road emergency; the assessment result generation unit 122 is used to make a judgment based on the weighted probability distribution of the category and risk level of the road emergency to obtain the assessment result of the road emergency.

[0054] The emergency instruction generation module 130 includes an evaluation result encoding and decoding unit 131 and an emergency instruction selection unit 132. The evaluation result encoding and decoding unit 131 is used to use an encoder and a decoder to calculate based on the evaluation result of the road emergency event to obtain a feature weighted context vector; the emergency instruction selection unit 132 is used to calculate the feature weighted context vector to generate an emergency instruction.

[0055] The system further includes an emergency instruction uploading module 140; the emergency instruction uploading module 140 is used to upload the generated emergency instruction to the cloud to remind the handling of road emergencies. After the emergency instruction is uploaded to the cloud, the backend can send the emergency instruction to relevant staff to inform them of the handling results of the road emergency.

[0056] It should be understood that the division of the various modules or units of the above system is only a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or physically separated. Moreover, these modules or units can be implemented in the form of software calling through processing elements; they can also be implemented in the form of hardware; some modules or units can be implemented in the form of software calling through processing elements, and some modules or units can be implemented in the form of hardware.

[0057] For example, the above modules or units may be one or more integrated circuits configured to implement the above methods, such as one or more application specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs). For another example, when a module is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call program code. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0058] like Figure 2 As shown, a flow chart of a method for automatically generating emergency instructions for handling road emergencies in an embodiment of the present application is shown. It mainly includes the following steps:

[0059] Step S11: collecting multimodal data of road emergencies in real time and preprocessing the multimodal data.

[0060] Specifically, basic data of current road emergencies, including traffic volume, vehicle speed, pedestrian flow, weather, etc., can be collected through sensors, sound collection equipment, and image collection equipment in intelligent roadside equipment, or information collection equipment equipped by personnel at the scene of road emergencies, drones, robot dogs, or robots. Cloud map data can also be intelligently linked to obtain location information of hospitals, police stations, traffic police, fire brigades, etc. around the current road emergency. The multimodal data finally collected includes text data, audio data, image data, video data, etc.

[0061] Among them, sensors include temperature sensors, humidity sensors, PM2.5 sensors, wind direction sensors, wind speed sensors, noise sensors, rainfall sensors, atmospheric pressure sensors, solar radiation sensors, and negative oxygen ion sensors, which can collect environmental data around road emergencies. Audio collection equipment Sound collection equipment includes microphones.

[0062] Image acquisition equipment includes but is not limited to: cameras, video cameras, camera modules integrated with optical systems or CCD chips, camera modules integrated with optical systems and CMOS chips, etc. Image acquisition equipment can be used to collect image data around road emergencies. Image data includes image data and video data. Image data can be basic information of illegal vehicles or illegal persons captured by snapshot cameras. Video data can be video images recorded in real time by surveillance cameras, including traffic conditions (such as pedestrian flow, vehicle flow, vehicle speed, vehicle violations, etc.) and on-site environment (such as extreme weather, dangerous behavior, etc.).

[0063] Preprocessing the multimodal data, specifically preprocessing the text data, audio data, image data and video data, the preprocessing method mainly includes the following:

[0064] (1) Preprocess the text data as follows: Figure 3 As shown in the figure, firstly, the original text data is segmented, and the original text is split into a sequence of words or subwords, using spaces, punctuation marks, etc. as separators. The text after word segmentation is cleaned, including removing special characters, stop words, and unifying capitalization, etc., and converted into a list of words or subwords. Then the text is input into the pre-trained word embedding model to be converted into a word embedding vector. Among them, the word embedding model includes TF-IDF, Word2Vec, etc.

[0065] Word embedding vectors refer to assigning a fixed-length vector representation to each word. This length can be set by yourself, such as 300, which is actually much smaller than the dictionary length (such as 10,000). After obtaining the word embedding vectors of each word, calculating the angle between two word vectors can be used as a measure of the relationship between them. The following table shows the relationship between each word vector and other word vectors.

[0066] Table 1 Relationship between two vectors

[0067] anarchism 0.5 0.1 -0.1 origin -0.5 0.3 0.9 as 0.3 -0.5 -0.3 a 0.7 0.2 -0.3 term 0.8 0.1 -0.1 of 0.4 -0.6 -0.1 abuse 0.7 0.1 -0.4

[0068] Furthermore, the formula for calculating the similarity between two word texts is as follows:

[0069]

[0070] Among them, A and B are the word vectors of two words respectively.

[0071] (2) Preprocess the audio data, convert the audio data into image data by calculating the Mel Frequency Cepstral Coefficient (MFCC), and then use the pre-trained model to obtain the audio feature sequence. The process of calculating the Mel Frequency Cepstral Coefficient (MFCC) is as follows: Figure 4 As shown, the specific steps include the following:

[0072] Step 1: After inputting the voice data, the voice data is pre-emphasized, framed, and windowed. Pre-emphasis uses a high-pass filter to amplify the high frequency and eliminate the effects of the vocal cords and lips during the utterance process to compensate for the high frequency part of the audio signal suppressed by the pronunciation system, and also to highlight the high-frequency resonance peak; framing is because if we directly perform Fourier transform on the audio signal, we can only get the spectrum of the entire audio segment, and there is no way to observe the change of the spectrum over time. For this reason, the audio needs to be cut into small segments, each of which is called a frame, which is framing; after framing, each frame is multiplied by a window function to smooth the signal, that is, windowing is performed, the purpose is to increase the continuity of both ends of the frame and reduce the subsequent operation of the spectrum leakage.

[0073] Step 2: Then, after FFT processing, the audio signal is converted from the time domain to the inverse frequency domain;

[0074] Step 3: Take the absolute value or square value to get the spectral line energy of the audio signal;

[0075] Step 4: Use Mel filtering (Mel filter analysis) to process the data and obtain the Mel spectrum through the Mel filter bank;

[0076] Step 5: Take the logarithm of the Mel spectrum;

[0077] Step 6: Use DCT to calculate the MFCC feature vector.

[0078] (3) Preprocessing the image data. The specific process includes adjusting, standardizing, and data augmentation to ensure that the image data have the same size and have features that are useful for model training.

[0079] First, the image data is resized, and can be resized to the same size as needed, such as 224x224 pixels.

[0080] The image data is then normalized to scale the pixel values ​​to a fixed range, usually [0, 1] or [-1, 1], which helps ensure that the pixel values ​​are in the same range, making the model easier to train. The normalization method is usually implemented by subtracting the mean (average pixel value) and dividing by the standard deviation, as shown in the following formula:

[0081]

[0082] Where μ is the mean of the image, X is the image matrix, and σ is the standard deviation of all pixel values.

[0083] Finally, the image data is enhanced. The data enhancement operations include rotation, mirror flipping, random cropping, brightness adjustment, etc. Figure 5 For example, operations such as rotation, mirror flipping, random cropping, and brightness adjustment are performed on image data.

[0084] (4) Preprocess the video data. Usually each video clip consists of multiple frames. The video data needs to be properly processed to handle time series and frame sequences. First, select a certain number of frames to extract from the video, or sample according to a specific time interval, and select one frame per second. For each sampled frame, image preprocessing is required, which includes operations such as image resizing, standardization, and data enhancement. This is achieved by arranging each frame on the timeline to form a video clip, thereby constructing a time series, and then extracting video sequence features through a pre-trained model.

[0085] Step S12: The pre-processed multimodal data is processed and judged using a neural network model and a hierarchical structure model to obtain the evaluation result of the road emergency. The specific steps are as follows: Figure 6 shown.

[0086] Step S121: using a neural network model to perform data characterization mapping on the preprocessed multimodal data to obtain a weighted probability distribution of the category and risk level of the road emergency event.

[0087] The neural network model includes a linear layer, a nonlinear layer, a classification layer and a fully connected layer. The specific process of using the neural network model to perform data representation mapping is as follows: Figure 7 As shown, including:

[0088] Step S1211: the linear layer performs a linear transformation on the preprocessed multimodal data to obtain a linear transformation result.

[0089] The preprocessed multimodal data is input into the linear layer of the neural network model and linearly transformed. Linear transformation is a mathematical operation that maps the values ​​in the input space to the output space by multiplying a weight matrix and / or adding operations. Linear transformation has the following characteristics: it can be expressed as matrix multiplication, the linear properties of the vector space are maintained, and the combination of multiple linear transformations is still linear. Linear transformation is expressed by the following formula:

[0090] y=W*x+b(Formula 3);

[0091] Among them, y is the linear transformation result, x is the input vector, W is the weight matrix of the linear transformation, and b is the offset vector.

[0092] Step S1212: The nonlinear layer performs a nonlinear transformation on the linear transformation result to obtain a high-dimensional dense vector.

[0093] The nonlinear layer receives the linear transformation result output by the linear layer, uses the activation function to perform nonlinear transformation, and then obtains a high-dimensional dense vector. Among them, nonlinear transformation is a mathematical transformation that performs nonlinear operations on the input data. It is usually used to introduce nonlinear characteristics to enable deep learning models to learn and execute more complex patterns and relationships. Nonlinear transformation usually uses activation functions, including Sigmoid activation function, ReLU activation function and Tanh activation function.

[0094] For example, taking text data as an example, the text data is preprocessed to obtain a word embedding vector, the word embedding vector is linearly transformed, and the linearly transformed word embedding vector is multiplied with the weight matrix to obtain a high-dimensional dense vector, usually 300 dimensions. Among them, the linear transformation process is as follows Figure 8 As shown, the linear transformation result is then transformed nonlinearly, such as using the ReLU activation function. The ReLU activation function has the characteristics of nonlinearity, sparsity, fast calculation speed and avoiding gradient disappearance. The formula of the ReLU activation function is as follows:

[0095] f(X)=max{0,X}(Formula 4);

[0096] Among them, f(X) represents the output of the ReLU activation function, X represents the input value, and max{0,X} represents the larger value between the input value and zero. The function image corresponding to the ReLU activation function is as follows Fig. 9 As shown, the high-dimensional dense vector finally obtained is as follows Fig.10 shown.

[0097] It should be noted that in neural networks, linear transformations and nonlinear transformations are usually applied alternately to form a combination of multiple levels to construct a part of a deep neural network. Linear and nonlinear transformations are used to transform the preprocessed multimodal data into a high-dimensional dense vector to perform enhanced representation processing on the preprocessed multimodal data, which can enhance the representation ability of the data.

[0098] Step S1213: The classification layer uses an activation function to perform classification calculation on the high-dimensional dense vector to obtain the probability distribution of the category and risk level of the road emergency.

[0099] It should be noted that in the classification layer of the neural network model, the high-dimensional dense vector is mapped to the probability distribution of the category and risk level space. Mapping to probability distribution means that the neural network model includes a classification layer, and the classification layer can map the results obtained after processing in the neural network to probability distribution, that is, after the input data is processed in the classification layer, the probability of each possible category is obtained, which is also the relative confidence of each category.

[0100] The activation function in this embodiment preferably uses the Softmax activation function, which is an activation function for multi-category classification. It converts the original score of each category (also called logits) into a probability value representing the probability distribution. The purpose of the Softmax activation function is to present the results of multi-classification in the form of probability.

[0101] The Softmax activation function exponentially and normalizes the raw scores of these categories so that the probability value of each category is between 0 and 1, and the sum of the probabilities of all categories is equal to 1. The final output vector represents the probability of each possible category. Through this process, the model can output the relative confidence of each category.

[0102] Specifically, the mathematical formula of the Softmax activation function is as follows:

[0103]

[0104] Among them, z i is the i-th element (raw score) of the input vector z. Indicates z i The index of represents the sum of all elements, and k is the dimension of vector z.

[0105] The specific implementation process of the Softmax activation function is as follows Fig.11 As shown, find the exponents of e for the three numbers respectively, and then find the proportion of the three numbers in the sum of the exponents. The one with the larger proportion is the prediction result.

[0106] Step S1214: the fully connected layer performs a weighted operation on the probability distribution of the category and risk level of the road emergency event to obtain a weighted probability distribution of the category and risk level of the road emergency event.

[0107] Specifically, the neural network model uses the weights of the fully connected layer to combine the probabilities of different categories and risk levels to obtain the final weighted probability. In this process, the neural network model takes into account the importance of different categories and risk levels, and associates them with each other through the adjustment of weights to generate the final category and risk level probabilities. This weighted operation allows the neural network model to better adapt to the needs of specific tasks and assign different weights to different categories and risk levels to obtain more accurate prediction and evaluation results.

[0108] It should be explained that the fully connected layer (Fully Connected Layer), also known as the densely connected layer, is a basic layer in a deep neural network. Fig.12 As shown in the figure, in the fully connected layer, each neuron is connected to each neuron in the previous layer, indicating that each input has an effect on each neuron in the current layer. The connection strength of the fully connected layer is determined by the weight and bias parameters, which are learned during training.

[0109] In this embodiment, after obtaining the weighted probability distribution of the categories and risk levels of road emergencies, it can be found that the categories of road emergencies can be divided into three parent categories, namely traffic accidents, crowd congestion, and natural disasters, each of which will have three subcategories of risk levels, namely severe, moderate, and minor.

[0110] Step S122: using a hierarchical model to make a judgment based on the weighted probability distribution of the category and risk level of the road emergency to obtain an assessment result of the road emergency.

[0111] Construct the road emergency category, and the hierarchy of risk level, and encode to obtain the hierarchical model, it should be noted that the purpose of encoding is to divide various road emergencies into different categories and risk levels, so as to better understand and handle them. In the present embodiment, first, a hierarchical model of dangerous events is established, and this structural model usually adopts a tree-like or hierarchical form to organize and display different categories and levels. The hierarchy can include multiple levels, starting from the large category of the high level, and gradually being subdivided into more specific subcategories and risk levels. Can start from the high-level "traffic accident", and then be divided into "serious", "medium", "slight" subcategories, forming different risk levels.

[0112] Next, code classification is performed. Each category and level requires a unique code for identification in data management and analysis. These codes can be numbers, letters, symbols, or any other form, as long as they are unique and identifiable. For example, a code of "1" can be assigned to "traffic accident" and a code of "1.1" can be assigned to "serious".

[0113] Finally, standardization is performed to ensure that the hierarchical model and coding system are consistent and standardized when organizing hazardous events. This can be achieved by defining clear rules and guidelines to ensure that all relevant personnel understand and use the same classification and coding methods, which is not limited in this embodiment.

[0114] It should be noted that after building the hierarchical model, the hierarchical model needs to be trained to obtain the optimal hierarchical model. Therefore, the loss function of each level in the hierarchical model needs to be determined and trained. Specifically, in the hierarchical classification problem, an appropriate loss function is selected for each level or level, and the model is trained using training data. Such methods are usually used for hierarchical classification tasks, where there is an obvious hierarchical relationship between categories.

[0115] In this embodiment, an appropriate loss function is selected for each layer or level, and a cross-entropy loss function is usually used to measure the difference between the output of the model and the true label. For hierarchical classification, a hierarchical cross-entropy loss function can be used, that is, the hierarchical information is combined with the loss function, and the relationship between different levels can be taken into account.

[0116] The cross entropy loss function, also known as the logarithmic loss function, is a common loss function used in machine learning and deep learning to measure the difference between two probability distributions. It is often used in classification problems, especially multi-class classification problems, to evaluate the degree of fit between the model's output probability distribution and the actual label. The formula for the multi-class cross entropy loss function is as follows:

[0117] L(y,p)=-∑ i y i log(p i )(Formula 6);

[0118] Among them, L(y,p) is the value of the loss function, y i is the one-hot encoding of the actual label, indicating that the sample belongs to the i-th category, p i is the predicted probability of the model, indicating the probability that the sample belongs to the i-th category. This loss function is used to measure the difference between the model's multi-category classification prediction and the actual label. It sums the cross entropy between the true label of each category and the predicted probability of the hierarchical model to obtain the total loss value. This ensures that the predicted probability distribution of the hierarchical model is as close as possible to the actual category distribution.

[0119] Finally, the classification judgment and risk level assessment results of road emergencies are output. For example, the output assessment result is "the risk level of traffic accidents is medium".

[0120] Step S13: Generate corresponding emergency instructions based on the evaluation results of the road emergency to complete the handling of the road emergency. Fig.13 As shown, the following steps are included:

[0121] Step S131: Based on the evaluation result of the road emergency, an encoder is used to calculate and obtain a feature context vector. The encoder includes an embedding layer, a position encoding layer, a first multi-head self-attention layer, and a first feedforward neural network layer. Fig.14 illustrate.

[0122] Step S1311: The embedding layer maps the evaluation result of the road emergency to obtain a continuous word embedding vector.

[0123] Specifically, an embedding layer in the encoder is used to map the category and risk level in the assessment result of the road emergency to a continuous word embedding vector, so that it is transformed from a discrete label to a continuous vector representation. This usually involves the use of an embedding layer, in which each category and level is associated with a unique embedding vector, which are learnable parameters of the model, and the learnable parameters are learned during the training process.

[0124] Step S1312: The position encoding layer adds position information to the word embedding vector to obtain a sequence representation result of the evaluation result.

[0125] The position encoding layer is introduced to consider the position information of the hierarchy of categories and risk levels. By adding the position information to the word embedding vector, the word embedding vector and the position encoding are combined to obtain a sequence representation result. The sequence representation result refers to the vector representation of the category and risk level and the position relationship result. That is, the encoder will process the evaluation results to capture the relationship and hierarchy between categories and risk levels, and generate the final output sequence representation result.

[0126] It should be noted that position encoding is usually achieved by combining position information with word embedding vectors of words or tags, so as to ensure that the positions of different elements in the sequence can be obtained. In this embodiment, relative position encoding is used. Relative position encoding acts on the self-attention mechanism to inform the distance between two elements. The result of relative position encoding is as follows: Fig.15 shown.

[0127] In this embodiment, the output evaluation result is "the risk level of a traffic accident is medium", and the order of relative position coding is "occurrence", "traffic accident", "risk level" is "medium".

[0128] Step S1313: The first multi-head self-attention layer calculates the sequence representation result to generate a context vector.

[0129] In the first multi-head self-attention layer, the importance distribution of different parts in the sequence is calculated by the multi-head self-attention mechanism, and then weighted and normalized to obtain the context vector. The multi-head self-attention mechanism is an important technology in deep learning. It calculates the importance distribution of different parts in the sequence, and then weights and normalizes them to generate a context vector. Specifically, the operation of the multi-head self-attention mechanism includes the following calculations:

[0130] Similarity calculation: The model calculates a similarity score between each position in the sequence and every other position. This is usually done using a dot product or other similarity measure.

[0131] Weight calculation: The similarity scores are subjected to a Softmax operation to obtain a weight distribution that represents the strength of association between each position and other positions.

[0132] Weighting and normalization: This step applies weights to all positions and weighted sums them to produce a context vector. This vector captures the importance of the current position relative to other positions, as well as the relationship between them.

[0133] The multi-head self-attention mechanism usually includes multiple sub-heads, each of which learns different correlation patterns to more comprehensively capture the information of different parts of the sequence. The final generated context vector is the concatenation or linear transformation of the context vectors generated by multiple sub-heads. The specific process of the multi-head self-attention mechanism is as follows Fig.16 As shown, the specific steps include:

[0134] Step 1: Get the original input sentence;

[0135] Step 2: Encode each word in the sentence to obtain matrix X; matrix X is the encoding result of the new sentence; matrix R is the direct output result of the previous layer;

[0136] Step 3: Divide it into 8 heads and multiply the matrix X or matrix R by each weight matrix;

[0137] Step 4: Calculate attention through the output query / key / value matrices;

[0138] Step 5: Concatenate all attention heads and multiply them by the weight matrix W o .

[0139] Among them, except for the first encoder, other encoders do not need to perform word embedding, and can directly use the output of the previous encoder as input (such as matrix R).

[0140] Step S1314: The first feedforward neural network layer performs nonlinear transformation and feature extraction on the context vector to obtain a feature context vector.

[0141] In the first feedforward neural network layer, the context vector is nonlinearly transformed and feature extracted by the feedforward neural network to better represent the features and relationships of the input sequence. The feedforward neural network contains multiple hidden layers, each of which includes multiple neurons. These neurons perform nonlinear transformations. The methods for performing nonlinear transformations include ReLU activation functions or Sigmoid activation functions. The results of these nonlinear mappings form feature representations.

[0142] Feedforward Neural Network (FNN), also known as Multi-Layer Perceptron (MLP), is a deep learning model. The key components of FNN include neurons and hierarchical structures. Fig.17 As shown, neurons are the basic units of the network, which are arranged in a hierarchical structure, usually including an input layer, a hidden layer, and an output layer. Each neuron is connected to all neurons in the previous layer and has weights, which are used to linearly transform the input data. Each neuron also has an activation function, which introduces nonlinear properties and enables the network to learn complex functional relationships.

[0143] Through the feedforward neural network, the context vector is nonlinearly transformed and feature extracted, with the aim of converting the context information of the input sequence into a more expressive feature representation to better capture the characteristics and relations of the input sequence and obtain the feature context vector.

[0144] Step S132: The feature context vector is processed in the decoder to obtain a feature weighted context vector.

[0145] The decoder includes a second multi-head self-attention layer, a second feedforward neural network layer, and a Fig.18 illustrate.

[0146] Step S1321: The second multi-head self-attention layer calculates the feature context vector to obtain a weighted context vector.

[0147] The second multi-head self-attention layer in the decoder calculates the degree of association between each position in the context vector through the multi-head self-attention mechanism, and performs weighted and normalized processing to obtain a weighted context vector. The decoder uses the context vector to better understand the input sequence and capture the contextual information of the input sequence when generating the output sequence. In this way, the multi-head self-attention mechanism allows the model to dynamically focus on different parts of the input sequence at different decoder positions, thereby improving the performance of the decoder in translation, text generation, and other sequence generation tasks.

[0148] Step S1322: The second feedforward neural network layer performs nonlinear transformation and feature extraction on the weighted context vector to obtain a feature weighted context vector.

[0149] The second feedforward neural network layer in the decoder performs nonlinear transformation and feature extraction on the weighted context vector through a feedforward neural network to increase the representation capability, with the goal of using contextual information to generate the output sequence. The weighted context vector is the output of the multi-head self-attention mechanism, which contains information about the importance of different positions in the input sequence. This vector is fed into the fully connected layers of the feedforward neural network, which perform nonlinear transformations. By learning weight parameters, these layers capture specific features or patterns in the context vector.

[0150] Through nonlinear transformation and feature extraction of feedforward neural networks, the weighted context vector is mapped to a more expressive representation. This helps the model better understand context information and better apply this information when generating output sequences. This process improves the representation ability of the model, enabling it to better understand and utilize context information to generate more accurate and coherent output sequences, namely feature weighted context vectors.

[0151] Step S133: Calculate the feature weighted context vector to generate a probability distribution of emergency instructions.

[0152] After obtaining the feature weighted context vector, a linear transformation and a normalized exponential function are used to calculate to generate the probability distribution of the next word or subword in the emergency instruction. In this embodiment, the Softmax activation function is also used to generate the probability of each word or character in the emergency instruction.

[0153] Step S134: A decision strategy is adopted to select and generate an optimal emergency instruction based on the probability distribution of the emergency instruction.

[0154] The next word or subword is selected from the probability distribution of the emergency instruction through a greedy strategy until the optimal emergency instruction is generated, and the emergency instruction is output to the instruction upload module. The greedy strategy is a decision-making strategy that is usually used for optimization problems and sequence generation tasks. The core idea is to select the option that currently looks optimal at each decision point without considering the impact of these choices on the optimal solution to the overall problem. In sequence generation tasks, the greedy strategy is usually used to gradually generate elements of the sequence, such as words, subwords, or symbols.

[0155] For example, in this embodiment, if the risk level of a traffic accident is medium, the emergency instruction generated may be to immediately call the police and contact medical personnel.

[0156] It should be emphasized that in order to achieve qualitative judgment of road emergencies, generative artificial intelligence technology can be used to use multimodal data, such as text data, audio data, image data and video data, to fully capture the different characteristics of sudden road dangerous events. Multimodal data can be represented and mapped to complete the classification of road emergencies and risk level assessment; at the same time, the qualitative judgment of sudden road dangerous events can be achieved by combining technologies such as hierarchical multi-label classification networks, which can effectively improve the timeliness and accuracy of judgment.

[0157] Furthermore, in order to realize the automatic generation of emergency instructions, generative artificial intelligence technology is used to utilize feedforward neural networks, multi-head self-attention mechanisms, encoder-decoder structure networks and other technologies to generate corresponding emergency instructions based on the assessment results of sudden road dangerous events, which can effectively improve the rationality and correctness of the issuance of emergency instructions.

[0158] like Fig.19 As shown, a schematic diagram of the structure of an electronic terminal in an embodiment of the present application is shown. The electronic terminal 1900 provided in this example includes: a memory 1901 and a processor 1902. The memory 1901 is used to store computer programs; the processor 1902 runs the computer program to implement the automatic generation method of emergency instructions for handling road emergencies.

[0159] Optionally, the number of the memories 1901 can be one or more, and the number of the processors 1902 can be one or more.

[0160] Optionally, the processor 1902 in the electronic terminal 1900 may be configured as follows: Figure 2 The steps described above load one or more instructions corresponding to the process of the application into the memory 1901, and the processor 1902 runs the application stored in the first memory 1901, thereby realizing various functions in the method for automatically generating emergency instructions for handling road emergencies.

[0161] Optionally, the memory 1901 may include but is not limited to high-speed random access memory and non-volatile memory. For example, one or more disk storage devices, flash memory devices or other non-volatile solid-state storage devices; the processor 1902 may include but is not limited to a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.

[0162] Optionally, the processor 1902 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0163] The present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method for automatically generating emergency instructions for handling road emergencies is implemented.

[0164] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to the computer program. The aforementioned computer program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, magnetic disk or optical disk and other media that can store program codes.

[0165] In the embodiments provided in the present application, the computer readable and writable storage medium may include a read-only memory, a random access memory, an EEPROM, a CD-ROM or other optical disk storage device, a disk storage device or other magnetic storage device, a flash memory, a USB flash drive, a mobile hard disk, or any other medium that can be used to store a desired program code in the form of an instruction or data structure and can be accessed by a computer. In addition, any connection can be appropriately referred to as a computer-readable medium. For example, if the instruction is sent from a website, a server or other remote source using a coaxial cable, an optical fiber cable, a twisted pair, a digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwaves, the coaxial cable, optical fiber cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwaves are included in the definition of the medium. However, it should be understood that computer readable and writable storage media and data storage media do not include connections, carriers, signals, or other temporary media, but are intended to be non-temporary, tangible storage media. Disk and disc, as used in this application, includes compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc where disks usually reproduce data magnetically, while discs reproduce data optically with lasers.

[0166] In summary, the present application provides. Therefore, the present application effectively overcomes various shortcomings in the prior art and has high industrial utilization value.

[0167] The above embodiments are merely illustrative of the principles and effects of the present application and are not intended to limit the present application. Anyone familiar with the technology may modify or change the above embodiments without violating the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by a person of ordinary skill in the art without departing from the spirit and technical ideas disclosed in the present application shall still be covered by the claims of the present application.

Claims

1. An automatic generation system of emergency instructions for handling road emergencies, It is characterized in that include: A multimodal data acquisition and processing module, used to acquire multimodal data of road emergencies in real time and pre-process the multimodal data; A road emergency event assessment module, used to process and judge the pre-processed multimodal data using a neural network model and a hierarchical structure model to obtain an assessment result of the road emergency event; The emergency instruction generation module is used to generate corresponding emergency instructions based on the evaluation result of the road emergency event.

2. According to claim 1, the automatic generation system of emergency instructions for handling road emergencies, It is characterized in that The system further comprises: Emergency instruction upload module: The emergency instruction upload module is used to upload the generated emergency instructions to the cloud to remind the handling of road emergencies.

3. A method for automatically generating emergency instructions for handling road emergencies, It is characterized in that The method applied to the automatic generation system of emergency instructions for handling road emergencies as claimed in claim 1 or 2 comprises: Collect multimodal data of road emergencies in real time and pre-process the multimodal data; The pre-processed multimodal data is processed and judged using a neural network model and a hierarchical structure model to obtain an assessment result of the road emergency; Based on the evaluation result of the road emergency event, a corresponding emergency instruction is generated to complete the handling of the road emergency event.

4. The method for automatically generating emergency instructions for handling road emergencies according to claim 3, It is characterized in that The specific process of using the neural network model and the hierarchical structure model to process and judge the pre-processed multimodal data to obtain the evaluation result of the road emergency includes: Using a neural network model to perform data characterization mapping on the preprocessed multimodal data to obtain a weighted probability distribution of the category and risk level of the road emergency event; A hierarchical structure model is used to make a judgment based on the weighted probability distribution of the category and risk level of the road emergency to obtain an assessment result of the road emergency.

5. The method for automatically generating emergency instructions for handling road emergencies according to claim 4, It is characterized in that The neural network model includes a linear layer, a nonlinear layer, a classification layer and a fully connected layer, wherein: The linear layer performs a linear transformation on the preprocessed multimodal data to obtain a linear transformation result; The nonlinear layer performs nonlinear transformation on the linear transformation result to obtain a high-dimensional dense vector; The classification layer uses an activation function to perform classification calculation on the high-dimensional dense vector to obtain the probability distribution of the category and risk level of the road emergency; The fully connected layer performs a weighted operation on the probability distribution of the category and risk level of the road emergency event to obtain a weighted probability distribution of the category and risk level of the road emergency event.

6. The method for automatically generating emergency instructions for handling road emergencies according to claim 3, It is characterized in that The specific process of generating corresponding emergency instructions based on the evaluation result of the road emergency includes: Based on the evaluation result of the road emergency, an encoder is used to calculate and obtain a feature context vector; The feature context vector is processed in a decoder to obtain a feature weighted context vector; Calculating the feature weighted context vector to generate a probability distribution of emergency instructions; A decision strategy is adopted to select and generate an optimal emergency instruction based on the probability distribution of the emergency instruction.

7. The method for automatically generating emergency instructions for handling road emergencies according to claim 6, It is characterized in that The encoder includes an embedding layer, a position encoding layer, a first multi-head self-attention layer, and a first feedforward neural network layer; wherein: The embedding layer maps the evaluation result of the road emergency to obtain a continuous word embedding vector; The position encoding layer adds position information to the word embedding vector to obtain a sequence representation result of the evaluation result; The first multi-head self-attention layer calculates the sequence representation result to generate a context vector; The first feedforward neural network layer performs nonlinear transformation and feature extraction on the context vector to obtain a feature context vector.

8. The method for automatically generating emergency instructions for handling road emergencies according to claim 6, It is characterized in that The decoder comprises a second multi-head self-attention layer and a second feed-forward neural network layer; wherein: The second multi-head self-attention layer calculates the feature context vector to obtain a weighted context vector; The second feedforward neural network layer performs nonlinear transformation and feature extraction on the weighted context vector to obtain a feature weighted context vector.

9. An electronic terminal, It is characterized in that include: Processor and memory; The memory is used to store computer programs; The processor is used to execute the computer program stored in the memory so that the electronic terminal executes the method for automatically generating emergency instructions for handling road emergencies as described in any one of claims 3 to 8.

10. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the method for automatically generating emergency instructions for handling road emergencies according to any one of claims 3 to 8 is implemented.