Bus full-life-cycle safety management method and system combined with big data

The method improves public vehicle management by using text embedding and neural networks to extract and combine features from public vehicle lifecycle data, addressing inaccuracies in existing evaluation methods and enhancing resource utilization.

CN119941480BActive Publication Date: 2025-07-15GUIYANG JINYANG CONSTR DATA SERVICE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510421869.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-15
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The existing bus standardized use evaluation methods are difficult to comprehensively and accurately capture the information in the archive record dimensions, resulting in the difficulty of single-dimensional non-standard vehicle use behavior being effectively identified, affecting the accuracy of bus standardized use evaluation.

Method used

By obtaining the historical record text of the bus life cycle archive, extracting text fragments of the archive record dimension and mapping them into text embedding vectors, using mobile ruler processing and vector integration technology, a bus usage characterization vector is generated, and the usage specification evaluation results are determined.

Benefits of technology

It improves the accuracy of evaluation of standardized use of buses, enhances the ability to identify single-dimensional irregular vehicle use behaviors, and provides an important basis for bus management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941480B_ABST
    Figure CN119941480B_ABST
Patent Text Reader

Abstract

The present application provides a bus full - life - cycle safety management method and system combined with big data. The method includes: obtaining the historical record text of the bus life - cycle file; extracting file text segments of multiple record dimensions from the historical record text, and mapping the extracted multiple file text segments into multiple text embedding vectors; using the first vector coverage range as the window size and the second vector coverage range as the window step, performing a sliding window process on each text embedding vector to obtain a first set of text sub - vectors corresponding to each text embedding vector; performing vector integration processing on each first set of text sub - vectors, and combining the multiple integrated text vectors obtained by integration to obtain a bus usage representation vector of the bus life - cycle file; determining the evaluation result of the usage specification of the bus life - cycle file through the bus usage representation vector. The present application can improve the accuracy of the evaluation of bus usage specifications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of text processing, and in particular, to a bus full-life-cycle safety management method and system combined with big data. Background Art

[0002] In the daily operations of enterprises and institutions, as an important means of transportation for official business, the standardized use of official buses is of extremely important significance. There are many problems and challenges in the standardized use of official buses. The phenomenon of private use of official buses occurs from time to time. Some personnel violate the regulations and use official buses for private affairs, resulting in serious waste of public resources. The problem of over-standard configuration is also relatively prominent. Some units do not configure official buses according to the specified standards, leading to unreasonable utilization of resources. There are also deficiencies in vehicle maintenance management. A considerable number of official buses are not maintained and serviced on time, resulting in poor vehicle conditions. At the same time, the vehicle repair and maintenance costs are relatively high, and the management method needs to be further optimized. The existing evaluation methods for the standardized use of official buses have certain limitations. Traditional evaluation methods often have difficulty in comprehensively and accurately capturing the information in the dimension of file records. The normative evaluation features of a single file dimension that are relatively concentrated in a short period are easily ignored in the global text embedding vector, making it difficult to effectively identify the non-standard vehicle use behavior in a single dimension, thereby affecting the accuracy of the evaluation of the standardized use of official buses. Therefore, a more effective technical solution is needed to improve the accuracy and reliability of the evaluation of the standardized use of official buses, so as to better strengthen the management of official buses, ensure that the use of official buses meets the specification requirements, and give full play to the role of official buses in official activities. Summary of the Invention

[0003] In view of this, at least one bus full-life-cycle safety management method and system combined with big data are provided in the embodiments of the present application. The technical solution of the present application is implemented as follows:

[0004] On the one hand, an embodiment of the present application provides a method for the full - life - cycle safety management of public vehicles combined with big data. The method includes: obtaining the historical archival record text of the public vehicle life - cycle file, where the public vehicle life - cycle file is a record set for the full - life - cycle management of the target public vehicle; extracting archival text segments of multiple archival record dimensions from the historical archival record text, and mapping the multiple extracted archival text segments into multiple text embedding vectors; where the multiple archival record dimensions include vehicle affairs file dimension, vehicle file dimension, driver file dimension, and violation file dimension; using the first vector coverage range as the extraction size and the second vector coverage range as the extraction step, performing a moving extraction process on each text embedding vector to obtain a first set of text sub - vectors corresponding to each text embedding vector; performing vector integration processing on each first set of text sub - vectors, and combining the multiple integrated text vectors obtained by integration to obtain the public vehicle usage characterization vector of the public vehicle life - cycle file; determining the usage specification evaluation result of the public vehicle life - cycle file through the public vehicle usage characterization vector.

[0005] On the other hand, the present application provides a computer system, including a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the program, it implements the steps in the above - mentioned method.

[0006] The method for the full - life - cycle safety management of public vehicles combined with big data provided by the present application obtains the historical archival record text of the public vehicle life - cycle file; extracts archival text segments of multiple archival record dimensions from the historical archival record text, and maps the multiple extracted archival text segments into multiple text embedding vectors; uses the first vector coverage range as the extraction size and the second vector coverage range as the extraction step, performs a moving extraction process on each text embedding vector to obtain a first set of text sub - vectors corresponding to each text embedding vector; performs vector integration processing on each first set of text sub - vectors, and combines the multiple integrated text vectors obtained by integration to obtain the public vehicle usage characterization vector of the public vehicle life - cycle file; determines the usage specification evaluation result of the public vehicle life - cycle file through the public vehicle usage characterization vector.

[0007] Through the above - mentioned solution, the present application adopts the technology of moving extraction processing to extract multiple extraction features (that is, the fusion result of text vectors of several attribute dimensions) from the text embedding vectors of each archival record dimension. In this way, a vector combination with richer information in each archival record dimension can be obtained. Then, it can prevent the normative evaluation features of a single archival dimension that are relatively concentrated in a short period of time from being ignored in the global text embedding vector, so as to increase the recognition ability of non - standard vehicle use in a single dimension and improve the accuracy of the public vehicle standard use evaluation.

[0008] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the technical solutions of this application. Description of the Drawings

[0009] Figure 1 It is a schematic diagram of the implementation process of a public vehicle full-life cycle safety management method combined with big data provided by an embodiment of this application.

[0010] Figure 2 It is a schematic diagram of the hardware entity of a computer system provided by an embodiment of this application. Detailed Embodiments

[0011] In order to make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be further elaborated in detail below in conjunction with the drawings and embodiments. The described embodiments should not be regarded as limitations of this application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.

[0012] An embodiment of this application provides a public vehicle full-life cycle safety management method combined with big data, and this method can be executed by a processor of a computer system. Among them, the computer system may refer to devices with data processing capabilities such as servers, laptop computers, tablet computers, desktop computers, mobile devices (such as mobile phones, portable video players, personal digital assistants), etc.

[0013] Figure 1 It is a schematic diagram of the implementation process of a public vehicle full-life cycle safety management method combined with big data provided by an embodiment of this application. As Figure 1 shown, this method includes:

[0014] Step 100: Obtain the historical file record text of the public vehicle life cycle file, where the public vehicle life cycle file is a record set for the full-life cycle management of the target public vehicle.

[0015] The public vehicle is a government vehicle, such as a vehicle used by enterprises and institutions for official activities. The full-life cycle management covers the entire process of the public vehicle from purchase, use, maintenance to scrapping, including information in the vehicle affairs file dimension, such as annual inspection, insurance, maintenance, etc.; information in the vehicle file dimension, such as materials related to the basic information of the vehicle; information in the driver file dimension, such as materials related to the basic information of the driver; and information in the violation file dimension, such as violation records, violators, penalty contents, etc. The historical file record text is a collection of these full-life cycle management information recorded in text form.

[0016] For example, for an official vehicle of an organization, its official vehicle life cycle file may include: vehicle file information such as invoices at the time of vehicle purchase, vehicle models and configuration information; vehicle operation file information such as annual inspection reports and insurance contracts for each year; driver file information such as driver's licenses and work resumes; and violation file information such as traffic violation tickets and handling results generated during vehicle use. These information are stored in the form of documents, reports, tables, etc., and need to be collected and organized into historical file record texts for subsequent processing.

[0017] To obtain these historical file record texts, the following methods can be adopted. If these file records are stored in paper documents, optical character recognition (OCR) technology can be used to convert the text in the paper documents into electronic text. OCR technology identifies and analyzes the characters in the image and converts them into a text format that can be processed by a computer. For example, for a paper annual inspection report, it can be scanned into an image file using a scanner, and then OCR software can be used to recognize the text in the image to finally obtain the annual inspection report in electronic text form.

[0018] If the file records are stored in electronic documents, these documents can be directly read. For structured electronic documents, such as tables in a database, SQL statements can be used for data query and extraction. For unstructured electronic documents, such as Word documents, PDF documents, etc., corresponding document parsing libraries can be used for text extraction. For example, for a Word document, the python-docx library in Python can be used to read the document content and convert it into text form.

[0019] Step 200: Extract archive text segments of multiple archive record dimensions from the historical archive record text, and map the multiple extracted archive text segments into multiple text embedding vectors; wherein, the multiple archive record dimensions include vehicle operation archive dimension, vehicle archive dimension, driver archive dimension, and violation archive dimension.

[0020] The multiple archive record dimensions include the vehicle operation archive dimension, covering information such as annual inspection, insurance, and maintenance; the vehicle archive dimension, involving materials related to vehicle basic information; the driver archive dimension, containing materials related to driver basic information; and the violation archive dimension, such as violation records, violators, penalty contents, etc.

[0021] For example, for the official vehicle of the aforementioned unit, the historical archive record text may be a comprehensive document containing various aspects of information. It is necessary to extract the archive text fragments of different archive record dimensions from it. For the vehicle affairs archive dimension, text fragments such as "The vehicle annual inspection was completed on October 15, 2024, and all indicators were qualified", "The vehicle insurance period is from January 1, 2024 to December 31, 2024, and the insurance amount is 500,000 yuan", and "On July 20, 2024, the vehicle was repaired due to engine failure, and the repair cost was 3,000 yuan" will be extracted. For the vehicle archive dimension, text such as "The vehicle brand is Volkswagen, the model is Passat, and the vehicle identification number is LSVCC2A46FN012345" will be extracted. For the driver archive dimension, content such as "The driver's name is Zhang San, the driver's license number is 1234567890, and the permitted driving type is C1" will be extracted. For the violation archive dimension, text fragments such as "On August 5, 2024, driver Li Si exceeded the speed limit on Zhongshan Road, was fined 200 yuan, and had 3 points deducted" will be extracted.

[0022] To extract these archive text fragments, a rule-based method can be adopted. By defining a series of rules, such as keyword matching, regular expressions, etc., to locate and extract the text fragments of different archive record dimensions. For example, for the annual inspection information in the vehicle affairs archive dimension, the rule can be defined as matching sentences containing keywords such as "annual inspection" and "yearly inspection". For the vehicle model information in the vehicle archive dimension, regular expressions can be used to match vehicle models in a specific format. Another method is a machine learning-based method, such as a named entity recognition (NER) model. A pre-trained NER model can be used to process the historical archive record text and identify the entity information of different archive record dimensions. For example, by using pre-trained models such as BERT and through fine-tuning training, it can be made to identify entities in different dimensions such as vehicle affairs archives, vehicle archives, driver archives, and violation archives.

[0023] After extracting the archive text fragments of multiple archive record dimensions, these archive text fragments are mapped into multiple text embedding vectors. Text embedding vectors are the representations of texts converted into vector spaces for subsequent calculations and analyses. Text embedding methods are, for example, word vector models such as Word2Vec and GloVe, as well as deep learning-based language models such as BERT and XLNet. Taking Word2Vec as an example, by training a neural network model, each word is mapped into a vector space with a fixed dimension. For an archive text fragment, each word in it can be converted into the corresponding word vector, and then these word vectors are combined into a text embedding vector through methods such as averaging and summing. Suppose an archive text fragment is "The vehicle annual inspection is qualified". First, the three words "vehicle", "annual inspection", and "qualified" are respectively converted into the corresponding word vectors , and then obtain the text embedding vector of this text segment by taking the average value .

[0024] For deep learning-based language models such as BERT, the entire text segment can be directly encoded to obtain a text embedding vector with a fixed dimension. Input the archival text segment into the pre-trained BERT model, and the model will output the text embedding vector corresponding to this text segment. By mapping archival text segments with multiple archival record dimensions into multiple text embedding vectors, the text information is converted into a vector form that can be processed by a computer, providing a basis for subsequent operations such as moving window processing and vector integration processing, which helps to analyze the usage of official vehicles more accurately and conduct standardized evaluations.

[0025] Step 300: Use the first vector coverage as the window size and the second vector coverage as the window stride to perform moving window processing on each text embedding vector to obtain the first set of text sub-vectors corresponding to each text embedding vector.

[0026] Moving window processing is a method for local feature extraction on sequence data. By setting the window size and the window stride, a window is slid over the text embedding vector, and the vector within the window is intercepted as a sub-vector each time, thus obtaining a series of sets of sub-vectors.

[0027] Taking the official vehicles of the unit mentioned above as an example, archival text segments with different archival record dimensions have been mapped into multiple text embedding vectors. Suppose a text embedding vector in the vehicle affairs archive dimension is , and each here is a numerical vector representing a point of this text segment in the vector space.

[0028] The first vector coverage, i.e., the window size, determines the length of the vector intercepted each time. Suppose the first vector coverage is set to 3 and the second vector coverage, i.e., the window stride, is set to 2. Starting from the starting position of the text embedding vector, intercept the vector with a window size of 3 to obtain the first sub-vector . Then move the window according to the window stride of 2 to intercept the second sub-vector , and continue to move the window to obtain , and so on until no more sub-vectors that meet the window size can be intercepted. Finally, for this text embedding vector in the vehicle affairs archive dimension, the obtained first set of text sub-vectors is .

[0029] In actual operation, moving window processing can be implemented by means of loop traversal. For each text embedding vector, starting from the first element of the vector, perform interception operations according to the window size and the window stride. Suppose the text embedding vector is If the sliding window size is k and the sliding window stride is s, the starting position i of the sub-vector starts from 0 and increases by s each time until i + k > n. The j-th sub-vector , where .

[0030] This method of sliding window processing can extract multiple sliding window features from the text embedding vectors of each archival record dimension by setting appropriate sliding window sizes and strides. These sliding window features are the fusion results of text vectors in several attribute dimensions, which can make the vector combination information of each archival record dimension more abundant. For example, in the dimension of train operation archives, vehicle maintenance information that is relatively concentrated within a short period may be ignored in the global text embedding vector. However, through sliding window processing, this part of the information can be extracted separately to form a valuable sub-vector, thereby increasing the recognition ability of non-standard vehicle use in a single dimension and improving the accuracy of the evaluation of the standardized use of official vehicles.

[0031] Step 400: Perform vector integration processing on each set of first text sub-vectors, and combine the multiple integrated text vectors obtained by integration to obtain the vehicle use representation vector of the official vehicle life cycle archive.

[0032] In step 400, perform vector integration processing on each set of first text sub-vectors, and combine the multiple integrated text vectors obtained by integration to obtain the vehicle use representation vector of the official vehicle life cycle archive. Vector integration processing is to merge and transform each set of first text sub-vectors corresponding to each archival record dimension to extract more representative features. The vehicle use representation vector is a vector representation that comprehensively reflects the use situation of official vehicles throughout their life cycle.

[0033] Taking the official vehicles of the unit mentioned above as an example, sliding window processing has been performed on the text embedding vectors of different archival record dimensions to obtain the set of first text sub-vectors corresponding to each dimension. For example, the set of first text sub-vectors in the dimension of train operation archives is , the set of first text sub-vectors in the dimension of vehicle archives is , the set of first text sub-vectors in the dimension of driver archives is , and the set of first text sub-vectors in the dimension of violation archives is .

[0034] Multiple methods can be used for vector integration processing. One method is to use neural network models such as Long Short-Term Memory (LSTM) or Gated Recurrent Unit (GRU). These models can process sequential data, model the first set of text sub-vectors, and capture the time series information therein. Taking LSTM as an example, each sub-vector in the first set of text sub-vectors is sequentially input into the LSTM model. The model updates the hidden state based on the current input and the hidden state at the previous moment, and finally outputs an integrated vector. Let the first set of text sub-vectors be , the hidden state update formula of the LSTM model is , where h t is the hidden state at time t, V t is the input sub-vector at time t, and h t-1 is the hidden state at the previous moment. After m moments of processing, the final hidden state h m is the integrated vector.

[0035] Another method is to use the attention mechanism. The attention mechanism can make the importance of different sub-vectors be taken into account when integrating vectors. Calculate the attention weights of each sub-vector, and then perform weighted summation on all sub-vectors according to these weights. Let the first set of text sub-vectors be , the attention weights be , then the integrated vector , where .

[0036] After completing the vector integration processing for each archival record dimension, multiple integrated text vectors will be obtained. For example, the integrated text vector for the train operation archive dimension is , the integrated text vector for the vehicle archive dimension is , the integrated text vector for the driver archive dimension is , and the integrated text vector for the violation archive dimension is . Next, these integrated text vectors are combined to obtain a bus usage representation vector. One way is to concatenate these vectors, that is, . Another method is to perform weighted combination according to the importance of each integrated text vector. The evaluation contribution coefficients of each integrated text vector can be determined according to information such as vehicle weighting levels, and then each integrated text vector is multiplied by the corresponding evaluation contribution coefficient and added together. Let the evaluation contribution coefficients be , then the bus usage representation vector .

[0037] By performing vector integration processing on each first set of text sub-vectors and combining the multiple integrated text vectors obtained, a bus usage representation vector that can comprehensively reflect the usage situation of the entire life cycle of the bus is obtained.

[0038] Step 500: Determine the evaluation result of the usage specification of the bus life cycle file based on the bus usage representation vector.

[0039] In step 500, the evaluation result of the usage specification of the bus life cycle file is determined based on the bus usage representation vector. This bus usage representation vector is a vector representation that comprehensively reflects the usage of the bus throughout its life cycle, containing information from multiple dimensions such as vehicle operation files, vehicle files, driver files, and violation files. To achieve this goal, a pre-trained model can be utilized, such as the second text processing network. This network is trained with a large number of bus life cycle file samples and can learn the mapping relationship between the bus usage specification and the bus usage representation vector. The obtained bus usage representation vector is input into the second text processing network, and the network performs a fully connected mapping on this vector. A fully connected mapping means that each neuron in the network is connected to all neurons in the previous layer. Through a series of linear transformations and non-linear activation functions, the input bus usage representation vector is converted into a probability distribution, namely the first evaluation probability distribution.

[0040] For example, assume that the evaluation results of the usage specification are divided into three cases: "standard", "basically standard", and "non-standard". The first evaluation probability distribution output by the second text processing network after the fully connected mapping may be [0.2, 0.3, 0.5], representing the probabilities that the bus usage situation is "standard", "basically standard", and "non-standard" respectively.

[0041] Determine the final evaluation result of the usage specification based on this first evaluation probability distribution. One method is to select the category with the highest probability as the evaluation result. In the above example, the probability of "non-standard" is the highest, which is 0.5. Therefore, it is determined that the evaluation result of the bus usage specification is "non-standard".

[0042] In specific implementation, the second text processing network can adopt a neural network structure such as a multi-layer perceptron (MLP). The MLP consists of an input layer, a hidden layer, and an output layer. The input layer receives the bus usage representation vector, the hidden layer performs non-linear transformations on the input through a series of neurons, and the output layer outputs the first evaluation probability distribution. Its calculation formula can be expressed as , where x is the bus usage representation vector, and are the weight matrix and bias vector of the i-th layer respectively, f is the activation function, such as ReLU, Sigmoid, etc., and y is the first evaluation probability distribution.

[0043] In this way, it is possible to accurately determine the evaluation result of the usage specification of the bus life cycle file by using the bus usage representation vector with the help of the second text processing network, providing an important basis for the management and supervision of the bus.

[0044] In one implementation, in step 400, vector integration processing is performed on each first text sub-vector set, and multiple integrated text vectors obtained by integration are combined to obtain a bus usage representation vector of the bus life cycle file, including:

[0045] Step 410: Determine the first text processing network corresponding to each first text sub-vector set;

[0046] Step 420: Load the first text sub-vector set corresponding to each text embedding vector into the corresponding first text processing network for vector integration processing to obtain an integrated text vector corresponding to each text embedding vector;

[0047] Step 430: Perform vector combination on multiple integrated text vectors to obtain a bus usage representation vector of the bus life cycle file.

[0048] In step 410, the first text processing network corresponding to each first text sub-vector set is determined. The first text sub-vector sets of different file record dimensions have different characteristics and information structures, and targeted processing networks are required for effective feature extraction and integration. Taking the previously mentioned unit official vehicles as an example, the first text sub-vector set of the vehicle affairs file dimension contains information such as vehicle annual inspection, insurance, and maintenance. These information have associations in time series and business logic; the first text sub-vector set of the vehicle file dimension is mainly the basic information of the vehicle, which is relatively static and structured; the first text sub-vector set of the driver file dimension involves information such as the driver's qualifications and driving habits; the first text sub-vector set of the violation file dimension records the violation situations of the vehicle, including violation time, location, type, etc. Appropriate first text processing networks will be allocated to each first text sub-vector set according to the characteristics of these different dimensions. These networks can be neural networks based on deep learning. For example, the long short-term memory network (LSTM) is suitable for processing the vehicle affairs file dimension with time series characteristics; the convolutional neural network (CNN) may be more suitable for feature extraction of structured information in the vehicle file dimension; the Transformer architecture can be used for the driver file and violation file dimensions to capture complex dependencies between information.

[0049] In step 420, the first text sub-vector set corresponding to each text embedding vector is loaded into the corresponding first text processing network for vector integration processing to obtain an integrated text vector corresponding to each text embedding vector. Taking the vehicle affairs file dimension as an example, assume that its first text sub-vector set contains multiple sub-vectors reflecting the vehicle status at different time points, such as etc. When using an LSTM network for processing, through its internal memory cells and gating mechanisms, the LSTM network can process these sub-vectors sequentially and capture long-term dependencies in the time series. The core formulas of LSTM include the input gate the forget gate the cell state update and the output gate where x t is the current input sub-vector, h t-1 is the hidden state at the previous moment, W is the weight matrix, b is the bias vector, is the Sigmoid function, is the element-wise multiplication. After a series of processes, finally an integrated vector is output, that is, the integrated text vector of the station operation file dimension. For the vehicle file dimension, if a CNN is used for processing, the CNN will slide the convolutional kernel over the first set of text sub-vectors to extract local features, and reduce the feature dimension through a pooling operation, and finally obtain the integrated text vector.

[0050] In step 430, vector combination is performed on multiple integrated text vectors to obtain the vehicle usage characterization vector of the public vehicle life cycle file. The integrated text vector of each file record dimension contains important information about the vehicle usage in that dimension, but the information of a single dimension is not sufficient to comprehensively reflect the overall vehicle usage situation. Therefore, it is necessary to combine the integrated text vectors of the station operation file dimension, the vehicle file dimension, the driver file dimension, and the violation file dimension. One method is the concatenation operation, that is, connecting these integrated text vectors in a certain order. Suppose the integrated text vector of the station operation file dimension is the integrated text vector of the vehicle file dimension is the integrated text vector of the driver file dimension is the integrated text vector of the violation file dimension is then the vehicle usage characterization vector . Another way is the weighted sum. According to the importance of different file record dimensions in the vehicle usage evaluation, corresponding weights are assigned to each integrated text vector, and then the weighted integrated text vectors are added together to obtain the vehicle usage characterization vector. Let the weights of the station operation file dimension, the vehicle file dimension, the driver file dimension, and the violation file dimension be and then the vehicle usage characterization vector .

[0051] Through steps S410 - S430, the conversion from the first set of text sub - vectors in different archival record dimensions to the vehicle - use characterization vector of the bus life - cycle archive is completed. This process fully considers the characteristics and importance of information in different dimensions, integrates the scattered information into a comprehensive vector representation, and lays a solid foundation for accurately evaluating the vehicle - use norms of the bus in the subsequent process. Different first - text processing networks and vector combination methods can be selected and optimized according to the actual situation to adapt to different bus management requirements and data characteristics.

[0052] In one implementation, in step S430, vector combination is performed on multiple integrated text vectors to obtain the vehicle - use characterization vector of the bus life - cycle archive, including:

[0053] Step S431: Determine the evaluation contribution coefficient of each integrated text vector; among them, the process of determining the evaluation contribution coefficient includes: obtaining the vehicle empowerment level; according to the correlation information between each archival record dimension and the vehicle empowerment level, determine the evaluation contribution coefficient of each integrated text vector.

[0054] Step S432: Perform weighted adjustment on each integrated text vector through the evaluation contribution coefficient to obtain multiple adjusted text vectors;

[0055] Step S433: Perform vector combination on multiple adjusted text vectors to obtain the vehicle - use characterization vector of the bus life - cycle archive.

[0056] First, obtain the vehicle empowerment level. The vehicle empowerment level contains information such as who uses the vehicle and the usage threshold, which reflects the importance and normative requirements of bus use. Taking the official vehicles of a unit as an example, the buses used by personnel at different levels of positions may have different empowerment levels. For the buses used by senior - level personnel, since their usage scenarios often involve important government affairs activities, the empowerment level is relatively high, and the requirements for vehicle maintenance, insurance, driver qualifications, etc. are also more stringent; while for the buses used by ordinary staff, the empowerment level is relatively low, and the usage norms and requirements may be relatively loose. The vehicle empowerment level information can be obtained from the database of the bus management system, and this information can be stored in a structured data form, such as a vehicle usage permission table, which records relevant information such as the positions of the users corresponding to each bus and the usage scope.

[0057] In step 4302, according to the correlation information between each file record dimension and the vehicle empowerment level, the evaluation contribution coefficient of each integrated text vector is determined. The evaluation contribution coefficient reflects the importance of each file record dimension in the evaluation of the official vehicle usage specification. For official vehicles with a higher empowerment level, the vehicle affairs file dimension (such as annual inspection, insurance, maintenance, etc. information) and the driver file dimension (such as materials related to the driver's basic information) are closely related to the normal use and safety guarantee of the vehicle. Therefore, the evaluation contribution coefficients of the integrated text vectors of these two dimensions may be relatively high. For official vehicles with a lower empowerment level, the violation file dimension (such as violation records, violators, penalty content, etc.) may account for a relatively large proportion in the evaluation because the daily use of such official vehicles focuses more on standardization. The evaluation contribution coefficient can be determined by establishing a correlation model. For example, a linear regression model can be used, with the vehicle empowerment level as the independent variable and the importance score of each file record dimension as the dependent variable. Through training with a large amount of historical data, the linear relationship between each file record dimension and the vehicle empowerment level is obtained, so as to calculate the evaluation contribution coefficient. Suppose the evaluation contribution coefficients of the vehicle affairs file dimension, vehicle file dimension, driver file dimension, and violation file dimension are respectively , and . For an official vehicle with a higher empowerment level, it may be ; while for an official vehicle with a lower empowerment level, it may be .

[0058] After determining the evaluation contribution coefficient, in step 431, it is clarified that each integrated text vector has its corresponding evaluation contribution coefficient, and these coefficients are determined according to the correlation between the vehicle empowerment level and the file record dimension in the previous steps. For example, the integrated text vector V of the vehicle affairs file dimension cs corresponds to the evaluation contribution coefficient , the integrated text vector V of the vehicle file dimension vehicle corresponds to the evaluation contribution coefficient , the integrated text vector V of the driver file dimension driver corresponds to the evaluation contribution coefficient , and the integrated text vector V of the violation file dimension violation corresponds to the evaluation contribution coefficient .

[0059] In step 432, each integrated text vector is weighted and adjusted through the evaluation contribution coefficient to obtain multiple adjusted text vectors. The purpose of weighted adjustment is to highlight the importance differences of different file record dimensions in the official vehicle usage evaluation. Multiply each integrated text vector by its corresponding evaluation contribution coefficient to obtain the adjusted text vector. Specifically, the adjusted text vector of the vehicle affairs file dimension is , and the adjusted text vector of the vehicle file dimension is , the adjusted text vector for the driver file dimension is , the adjusted text vector for the violation file dimension is . Taking the public bus with a higher empowerment level as an example, the integrated text vector V of the vehicle affairs file dimension cs after being multiplied by a higher evaluation contribution coefficient , its influence in the subsequent combination will be relatively greater because the vehicle affairs file is crucial for the normal operation and safety guarantee of high-level public buses.

[0060] In step 433, multiple adjusted text vectors are combined vectorially to obtain the vehicle usage representation vector of the public bus life cycle file. Vector combination is to integrate the information after weighted adjustment of each file record dimension to form a comprehensive vector representation to comprehensively reflect the usage situation of the public bus. The combination method adopted is to add multiple adjusted text vectors, and the formula is . In this way, the information of different file record dimensions is effectively fused according to their importance. For example, for a public bus with a higher empowerment level, the adjusted text vectors of the vehicle affairs file dimension and the driver file dimension have a greater impact on the final vehicle usage representation vector during the addition process, making the vector better reflect the requirements of high-level public buses in terms of vehicle affairs management and driver qualifications; while for a public bus with a lower empowerment level, the adjusted text vector of the violation file dimension occupies a more important position in the vehicle usage representation vector due to its higher evaluation contribution coefficient, highlighting the importance of the usage standardization of such public buses.

[0061] In actual operation, matrix operations can be used to efficiently complete these steps. For the evaluation contribution coefficient, it can be stored as a coefficient vector , the integrated text vector can be stored as a matrix , then the vehicle usage representation vector . Such a matrix operation method can make full use of the parallel computing power of the computer to improve the processing efficiency.

[0062] Through steps 431 - S433, the influence of the vehicle empowerment level on different file record dimensions is fully considered, the integrated text vectors are reasonably weighted and combined, and the vehicle usage representation vector that can accurately reflect the usage situation of the public bus is obtained.

[0063] In one implementation scheme, in step 200, after extracting the file text fragments of multiple file record dimensions from the historical file record text and mapping the obtained multiple file text fragments into multiple text embedding vectors, the method further includes:

[0064] Step 210: Use the coverage range of the third vector as the window size and the coverage range of the fourth vector as the window stride, and perform a sliding window operation on each text embedding vector to obtain a set of second text sub-vectors corresponding to each text embedding vector. Based on this, as another implementation of step 400, that is, perform vector integration processing on each set of first text sub-vectors, and combine the multiple integrated text vectors obtained by integration to obtain the vehicle usage representation vector of the bus life cycle file, including:

[0065] Step 401: Perform vector integration processing on each set of first text sub-vectors, and combine the multiple first integrated sub-text vectors obtained by integration to obtain a first vehicle usage sub-representation vector. Perform vector integration processing on each set of second text sub-vectors, and combine the multiple second integrated sub-text vectors obtained by integration to obtain a second vehicle usage sub-representation vector;

[0066] Step 402: Combine the first vehicle usage sub-representation vector and the second vehicle usage sub-representation vector to obtain the vehicle usage representation vector of the bus life cycle file.

[0067] In step 210, after extracting the historical archive record text into multiple archive text segments of archive record dimensions and mapping these segments into multiple text embedding vectors, a sliding window operation is performed on each text embedding vector again. This time, the coverage range of the third vector is used as the window size and the coverage range of the fourth vector is used as the window stride, so as to obtain a set of second text sub-vectors corresponding to each text embedding vector. It can be understood that the coverage ranges of each vector in the embodiments of the present application can be preset according to actual needs, and are not specifically limited. The purpose of step 210 is to extract features from the text embedding vectors at different scales and granularities to capture more information at different levels. Taking the unit official vehicle mentioned above as an example, assume that a text embedding vector in the vehicle affairs archive dimension contains vector representations of information such as vehicle annual inspection, insurance, and maintenance. The first sliding window operation (corresponding to step 300) may extract some features at a specific scale, while in step 210, different window sizes and strides are used for reprocessing. For example, when the window size is larger in the first processing, some more macroscopic information can be extracted, while in this processing, a smaller window size and stride are used, and more detailed local features can be mined, such as more accurate representations of the specific time and cost of a vehicle repair in the vector. The method for implementing this sliding window operation is similar to step 300, and the second text sub-vector set can be obtained by traversing the text embedding vector in a loop and performing intercepting operations according to the set window size and stride.

[0068] In step 401, vector integration processing is performed on each set of first text sub-vectors. The multiple first integrated sub-text vectors obtained through integration are combined to obtain a first bus usage sub-representation vector. At the same time, vector integration processing is performed on each set of second text sub-vectors, and the multiple second integrated sub-text vectors obtained through integration are combined to obtain a second bus usage sub-representation vector. For the vector integration processing of the set of first text sub-vectors, methods such as long short-term memory network (LSTM), gated recurrent unit (GRU), or attention mechanism mentioned above can be used. Taking LSTM as an example, for the set of first text sub-vectors in the vehicle affairs file dimension, the LSTM network will process the input sub-vector sequence according to its internal memory unit and gating mechanism, and output an integrated vector, that is, the first integrated sub-text vector. The first integrated sub-text vectors in multiple file record dimensions are combined by means such as concatenation or weighted summation to obtain a first bus usage sub-representation vector. Similarly, for the set of second text sub-vectors, the same or different vector integration methods are used for processing to obtain second integrated sub-text vectors and combine them into a second bus usage sub-representation vector. Suppose the first integrated sub-text vectors in the vehicle affairs file dimension, vehicle file dimension, driver file dimension, and violation file dimension are respectively , and a first bus usage sub-representation vector is obtained by weighted summation combination , where are the corresponding weight coefficients. Similarly, for the second integrated sub-text vectors , a second bus usage sub-representation vector is combined .

[0069] In step 402, the first bus usage sub-representation vector and the second bus usage sub-representation vector are combined to obtain a bus usage representation vector of the bus life cycle file. This combination process can further synthesize feature information at different scales and levels, making the final bus usage representation vector more comprehensive and accurate. The combination method can be concatenation or weighted summation.

[0070] Through step 210 and steps 401 - 402, feature extraction and integration are performed on the text information of the bus life cycle file from different scales and levels, and finally a more comprehensive and accurate bus usage representation vector is generated. This vector can better reflect the usage of the bus during its entire life cycle, provide a basis for the subsequent evaluation of bus usage specifications, help improve the accuracy and reliability of the evaluation, and meet the actual needs of bus management.

[0071] In one implementation scheme, in step 500, the usage specification evaluation result of the bus life cycle file is determined through the bus usage representation vector, including:

[0072] Step 510: Load the bus usage characterization vector into the second text processing network for fully connected mapping to obtain the first evaluation probability distribution;

[0073] Step 520: Determine the usage specification evaluation result of the bus life cycle file according to the first evaluation probability distribution.

[0074] In step 510, the bus usage characterization vector is loaded into the second text processing network for fully connected mapping to obtain the first evaluation probability distribution. The second text processing network is a pre-trained neural network model, such as a multi-layer perceptron (MLP). The multi-layer perceptron consists of an input layer, a hidden layer, and an output layer. The input layer receives the bus usage characterization vector. The hidden layer contains multiple neurons that process the input through a series of linear transformations and non-linear activation functions. The output layer outputs the first evaluation probability distribution.

[0075] Suppose the bus usage characterization vector is V bur , and its dimension is n. The input layer of the second text processing network has n neurons, corresponding to the dimension of the bus usage characterization vector. The number of neurons in the hidden layer can be adjusted according to the actual situation. Suppose the hidden layer has m neurons. The linear transformation from the input layer to the hidden layer can be expressed as , where W1 is an m×n weight matrix and b_1 is an m-dimensional bias vector. Then, process Z1 through a non-linear activation function f (such as the ReLU function, f(x)=max(0,x)) to obtain the output H=f(Z1) of the hidden layer.

[0076] The linear transformation from the hidden layer to the output layer is , where W2 is a k×m weight matrix (k is the number of neurons in the output layer, corresponding to the number of categories of the evaluation result), and b2 is a k-dimensional bias vector. Finally, convert Z2 to a probability distribution through the Softmax function, that is, the first evaluation probability distribution P = Softmax(Z2). The formula of the Softmax function is , where z i is the i-th element of Z2.

[0077] Suppose the bus usage specification evaluation results are divided into three cases: "standard", "basically standard", and "non-standard". Then the output layer has 3 neurons, and the first evaluation probability distribution P = [p1, p2, p3], which respectively represent the probabilities that the bus usage situation is "standard", "basically standard", and "non-standard", and p1 + p2 + p3 = 1.

[0078] In step 520, the usage specification evaluation result of the bus life cycle file is determined according to the first evaluation probability distribution. One method is to select the category with the highest probability as the evaluation result. For example, if the first evaluation probability distribution is [0.2, 0.3, 0.5], then the probability of "non-standard" is the highest, and the usage specification evaluation result of this bus is determined to be "non-standard".

[0079] To implement this process, the Argmax function can be used, that is, Result = argmax(P), where Result is the index of the evaluation result. If Result = 0, the evaluation result is "standard"; if Result = 1, the evaluation result is "basically standard"; if Result = 2, the evaluation result is "non-standard".

[0080] In practical applications, the training of the second text processing network is based on a large number of bus life cycle file samples. By adjusting the weight matrices W1, W2 and the bias vectors b1, b2, the network can accurately map the bus usage feature vectors to the corresponding usage specification evaluation results. The loss function (such as the cross-entropy loss function) is used in the training process to measure the difference between the prediction result and the true label, and the parameters of the network are updated through the backpropagation algorithm.

[0081] In one implementation scheme, multiple first text processing networks and the second text processing network are obtained through parallel training with bus life cycle file samples. The multiple first text processing networks and the second text processing network are trained through the following process:

[0082] Step 10: Obtain the file record training text, which includes the full-cycle file record training text of multiple bus training examples and the supervision label of the usage specification evaluation result of each bus training example;

[0083] Step 20: Extract file training text segments of multiple file record dimensions from the full-cycle file record training text, and map the extracted multiple file training text segments into multiple training text embedding vectors;

[0084] Step 30: Use the first vector coverage range as the sliding window size and the second vector coverage range as the sliding window step, and perform sliding window processing on each training text embedding vector to obtain the first training text sub-vector set corresponding to each training text embedding vector;

[0085] Step 40: Load the first training text sub-vector set corresponding to each training text embedding vector into the corresponding first text processing network for vector integration processing, and combine the training integration text vectors corresponding to each training text embedding vector obtained by integration to obtain the bus example usage feature vector;

[0086] Step 50: Load the bus sample usage representation vector into the second text processing network for full connection mapping to obtain a second evaluation probability distribution, and determine the mapping error value through the second evaluation probability distribution and the corresponding usage specification evaluation result supervision mark;

[0087] Step 60: Convergence training is performed on the network parameters of the plurality of first text processing networks and the second text processing networks by mapping the error values.

[0088] In step 10, the archive record training text is obtained, which includes the full-cycle archive record training text of multiple bus training examples and the supervision mark of the use specification evaluation result of each bus training example. The bus training example is a representative sample selected from the actual use data of a large number of buses. The full-cycle archive record training text covers various information from the purchase, use to scrapping of the bus, such as annual inspection, insurance, and maintenance records in the vehicle archive dimension, basic vehicle information in the vehicle archive dimension, driver qualifications and driving records in the driver archive dimension, and violations in the violation archive dimension. The supervision mark of the use specification evaluation result is a clear mark of whether the use of each bus training example is standardized, for example, it can be represented by labels such as "standard", "basic standard", and "non-standard". Taking the official vehicles of the unit as an example, the full-cycle archive records of 1,000 buses may be collected from the database as training texts, and the use specification evaluation results are marked for each bus, such as the first vehicle is marked as "standard", the second vehicle is marked as "non-standard", etc.

[0089] In step 20, multiple archive training text segments of multiple archive record dimensions are extracted from the full-cycle archive record training text, and the multiple extracted archive training text segments are mapped into multiple training text embedding vectors. A rule-based method or a machine learning method can be used to extract the archive training text segments. The rule-based method locates text segments of different archive record dimensions by defining keyword matching rules or regular expressions. For example, sentences containing keywords such as "annual review" and "insurance" are matched as text segments of the vehicle affairs archive dimension. Machine learning methods such as named entity recognition (NER) models can learn entity information of different archive record dimensions through training, so as to extract text segments more accurately. When mapping the archive training text segments into training text embedding vectors, a word vector model (such as Word2Vec, GloVe) or a deep learning-based language model (such as BERT, XLNet) can be used. Taking Word2Vec as an example, it maps each word into a vector space of a fixed dimension. For an archive training text segment, each word in it is converted into a corresponding word vector, and then these word vectors are combined into a training text embedding vector by means of averaging or summing. Suppose an archive training text segment is "The vehicle annual review is qualified". The three words "vehicle", "annual review", and "qualified" are respectively converted into word vectors v1, v2, and v3, and then the training text embedding vector V = (v1 + v2 + v3) / 3 is calculated.

[0090] In step 30, the first vector coverage range is used as the extraction size, and the second vector coverage range is used as the extraction stride. Each training text embedding vector is processed by moving extraction to obtain a set of first training text sub-vectors corresponding to each training text embedding vector. The moving extraction process is a method for local feature extraction on sequence data. By setting the extraction size and extraction stride, a sliding window is used on the training text embedding vector, and the vector within the window is intercepted each time as a sub-vector. An example can refer to the aforementioned step 300 and will not be elaborated here.

[0091] In step 40, the first training text sub-vector set corresponding to each training text embedding vector is loaded into the corresponding first text processing network for vector integration processing, and the training integrated text vectors corresponding to each training text embedding vector obtained by integration are combined to obtain a bus example usage representation vector. The first text processing network can be a long short-term memory network (LSTM), a gated recurrent unit (GRU), or a Transformer architecture, etc. Taking LSTM as an example, it can handle long-term dependencies in sequential data through internal memory units and gating mechanisms. After being processed by the LSTM network, a final integrated vector, that is, the training integrated text vector, is output. The training integrated text vectors of multiple archival record dimensions are combined, such as by concatenation or weighted summation, to obtain the bus example usage representation vector.

[0092] In step 50, the bus example usage representation vector is loaded into the second text processing network for fully connected mapping to obtain the second evaluation probability distribution, and the mapping error value is determined through the second evaluation probability distribution and the corresponding usage specification evaluation result supervision label. The second text processing network can be a multi-layer perceptron (MLP), which consists of an input layer, a hidden layer, and an output layer. The input layer receives the bus example usage representation vector, the hidden layer processes the input through a series of linear transformations and non-linear activation functions, and the output layer outputs the second evaluation probability distribution. Assuming that the bus usage specification evaluation results are divided into three cases: "specification", "basically specification", and "non-specification", the output layer has 3 neurons, and the second evaluation probability distribution P = [p1, p2, p3] represents the probabilities of the bus usage being "specification", "basically specification", and "non-specification" respectively. The cross-entropy loss function is used to measure the difference between the second evaluation probability distribution and the usage specification evaluation result supervision label. The formula of the cross-entropy loss function is , where is the i-th element of the usage specification evaluation result supervision label (if it is "specification", then ; if it is "basically specification", then ; if it is "non-specification", then ), p i is the i-th element of the second evaluation probability distribution, and k is the number of categories of evaluation results. This loss value is the mapping error value, which reflects the degree of difference between the network prediction result and the true label.

[0093] In step 60, the network parameter variables of multiple first text processing networks and second text processing networks are convergently trained through mapping error values. The backpropagation algorithm is used to calculate the gradients of the mapping error values with respect to the network parameter variables (such as weight matrices and bias vectors), and then the parameter variables are updated according to the gradients. Feasible optimization algorithms include Stochastic Gradient Descent (SGD), Adagrad, Adadelta, Adam, etc. Taking the stochastic gradient descent algorithm as an example, its update formula is , where represents the network parameter variable (such as weight matrix W or bias vector b), is the learning rate, which controls the step size of parameter variable update, is the gradient of the mapping error value with respect to the parameter variable. Through multiple iterative trainings, the network parameter variables are continuously adjusted to gradually reduce the mapping error value until the convergence condition is reached (such as the error value is less than a certain threshold or the number of iterations reaches the upper limit). At this time, the performance of the network reaches the optimal, and it can accurately extract features from the bus life cycle archives and evaluate the bus usage specifications.

[0094] In one implementation, in step 20, after extracting archive training text segments of multiple archive record dimensions from the full-cycle archive-recorded training text and mapping the extracted multiple archive training text segments into multiple training text embedding vectors, the method further includes:

[0095] Step 201: Using the third vector coverage range as the ruler-taking size and the fourth vector coverage range as the ruler-taking step, perform a moving ruler-taking process on each training text embedding vector to obtain a set of second training text sub-vectors corresponding to each training text embedding vector; based on this, in step 40, load the set of first training text sub-vectors corresponding to each training text embedding vector into the corresponding first text processing network for vector integration processing, and combine the training integrated text vectors corresponding to each integrated training text embedding vector to obtain a bus sample usage representation vector, including:

[0096] Step 41: Load the set of first training text sub-vectors corresponding to each training text embedding vector into the corresponding first text processing network for vector integration processing, and combine the first training integrated sub-text vectors corresponding to each integrated training text embedding vector to obtain a first training bus usage sub-representation vector;

[0097] Step 42: Load the set of second training text sub-vectors corresponding to each training text embedding vector into the corresponding first text processing network for vector integration processing, and combine the second training integrated sub-text vectors corresponding to each integrated training text embedding vector to obtain a second training bus usage sub-representation vector;

[0098] Step 43: Combine the first training bus usage sub-representation vectors with the second training bus usage sub-representation vectors to obtain the bus example usage representation vectors of the bus training examples.

[0099] Step 20 completed extracting the archival training text fragments of multiple archival record dimensions from the full-cycle archival record training text and mapping them into multiple training text embedding vectors. On this basis, in Step 201, the third vector coverage range is used as the window size and the fourth vector coverage range is used as the window stride to perform a sliding window operation on each training text embedding vector to obtain the second set of training text sub-vectors corresponding to each training text embedding vector. The purpose of this step is to extract features from the training text embedding vectors at different scales and granularities. Taking the archival records of unit official vehicles mentioned before as an example, assume that a training text embedding vector of the vehicle affairs archival dimension contains the annual inspection, insurance, and maintenance information of the vehicle over the years. In Step 30, a sliding window operation has been performed according to the first vector coverage range and the second vector coverage range to obtain the first set of training text sub-vectors. In Step 201, different window sizes and strides are used for reprocessing. For example, when processing for the first time, the window size is larger, and features about the overall vehicle affairs situation of a certain year of the vehicle may be extracted; while this time, smaller window sizes and strides are used to be able to dig out more detailed information, such as the detailed cost of a specific repair and the more accurate representation of the repair items in the vector. By looping through the training text embedding vectors and performing the interception operation according to the set third vector coverage range and fourth vector coverage range, the second set of training text sub-vectors is obtained.

[0100] Based on the second set of training text sub-vectors obtained in Step 201, in Step 41, the first set of training text sub-vectors corresponding to each training text embedding vector is loaded into the corresponding first text processing network for vector integration processing, and the first training integrated sub-text vectors corresponding to each training text embedding vector obtained by integration are combined to obtain the first training bus usage sub-representation vectors. The first text processing network can be a long short-term memory network (LSTM), a gated recurrent unit (GRU), or a Transformer architecture, etc. After a series of processing, an integrated vector, that is, the first training integrated sub-text vector, is output. For multiple archival record dimensions, such as the vehicle affairs archival dimension, the vehicle archival dimension, the driver archival dimension, and the violation archival dimension, the first training integrated sub-text vectors of each dimension are combined. The combination method can be concatenation or weighted summation.

[0101] In step 42, the set of second training text sub-vectors corresponding to each training text embedding vector is loaded into the corresponding first text processing network for vector integration processing, and the second training integrated sub-text vectors corresponding to each training text embedding vector obtained through integration are combined to obtain the second training bus usage sub-representation vector. Similar to step 41, the same or different first text processing networks are used to process the set of second training text sub-vectors. Since the set of second training text sub-vectors is obtained from different window-slicing scales, it contains feature information at different levels from the set of first training text sub-vectors. For example, for the vehicle file dimension, the set of second training text sub-vectors may extract more detailed features regarding vehicle configuration details.

[0102] In step 43, the first training bus usage sub-representation vector and the second training bus usage sub-representation vector are combined to obtain the bus example usage representation vector of the bus training example. This step synthesizes feature information at different scales and levels, enabling the final bus example usage representation vector to more comprehensively and accurately reflect the usage of the bus. The combination method can also be concatenation or weighted summation.

[0103] Through step 201 and steps 41 - S43, feature extraction and integration are performed on the training text embedding vectors of the bus life cycle file from different scales and levels, and finally a more comprehensive and accurate bus example usage representation vector is generated. This vector can better reflect the usage of the bus throughout its life cycle, providing a more powerful basis for subsequent evaluation of bus usage norms, helping to improve the accuracy and reliability of the evaluation, and meeting the actual needs of bus management.

[0104] In one implementation, each first text processing network includes a text pre-coding component and a vector integration component. At this time, as another implementation of step 40, that is, the set of first training text sub-vectors corresponding to each training text embedding vector is loaded into the corresponding first text processing network for vector integration processing, and the training integrated text vectors corresponding to each training text embedding vector obtained through integration are combined to obtain the bus example usage representation vector, including:

[0105] Step 4010: Perform text pre-coding on the set of first training text sub-vectors corresponding to each training text embedding vector through the text pre-coding component to obtain the first training pre-coded vector;

[0106] Step 4020: Perform vector integration processing on the first training pre-coded vector corresponding to each training text embedding vector through the vector integration component to obtain the training integrated text vector corresponding to each training text embedding vector;

[0107] Step 4030: Combine the training integrated text vectors corresponding to each training text embedding vector to obtain a bus example usage representation vector.

[0108] In step 4010, the text precoding component performs text precoding on the first training text sub-vector set corresponding to each training text embedding vector to obtain the first training precoded vector. The purpose of text precoding is to map the first training text sub-vector set to a latent space dimension that is more suitable for subsequent processing. Taking the file records of unit official vehicles mentioned above as an example, assume that the first training text sub-vector set in the vehicle affairs file dimension contains vector representations of relevant information such as vehicle annual inspection, insurance, and maintenance. The text precoding component can use linear projection to project each sub-vector in the first training text sub-vector set into the latent space dimension of the Transformer. The formula for linear projection can be expressed as V pe = W × V sv + b, where V sv is the sub-vector in the first training text sub-vector set, W is the weight matrix, b is the bias vector, and V pe is the first training precoded vector. Through this linear transformation, the original sub-vectors can be converted into precoded vectors with richer feature representations, which are convenient for the subsequent vector integration component to process.

[0109] In step 4020, the vector integration component performs vector integration processing on the first training precoded vector corresponding to each training text embedding vector to obtain the training integrated text vector corresponding to each training text embedding vector. The role of the vector integration component is to further process and integrate the precoded vectors to extract more representative features. The vector integration component can adopt various architectures, such as the multi-head attention mechanism in the Transformer architecture. The multi-head attention mechanism enables the model to focus on different parts of the input in different representation sub-spaces, thereby capturing the relationships between input vectors more comprehensively. The formula for the multi-head attention mechanism is , where , Q, K, and V are the query, key, and value matrices respectively, and are learnable weight matrices. Taking the first training precoded vector as the input, calculate the attention weights of different parts through the multi-head attention mechanism, and then perform weighted summation on the vectors to obtain the integrated vector, that is, the training integrated text vector.

[0110] In step 4030, the training integrated text vectors corresponding to each training text embedding vector are combined to obtain a bus example usage representation vector. The training integrated text vectors of different file record dimensions contain important information about bus usage in their respective dimensions, and these information need to be integrated. The combination method can be concatenation or weighted summation.

[0111] Through steps 4010 - S4030, using the text pre - encoding component and vector integration component in the first text processing network, the first training text sub - vector sets corresponding to each training text embedding vector are effectively processed and integrated. The text pre - encoding component projects the original sub - vectors into a latent space dimension more suitable for processing, and the vector integration component further extracts and integrates features, and finally combines the training integrated text vectors of different dimensions into a bus example usage representation vector. This process makes full use of the structure and function of the first text processing network, providing a more accurate and representative input vector for the subsequent bus usage specification evaluation.

[0112] In one implementation scheme, each first text processing network further includes a position embedding component. In step 4010, after obtaining the first training pre - encoding vectors by performing text pre - encoding on the first training text sub - vector sets corresponding to each training text embedding vector through the text pre - encoding component, the method further includes:

[0113] Step 4011: Perform position embedding on the first training text sub - vector sets corresponding to each training text embedding vector through the position embedding component to obtain a set of first training text position sub - vectors. Based on this, in step 4020, perform vector integration processing on the first training pre - encoding vectors corresponding to each training text embedding vector through the vector integration component to obtain the training integrated text vectors corresponding to each training text embedding vector, including:

[0114] Step 4021: Perform vector integration processing on the first training text position sub - vector sets corresponding to each training text embedding vector through the vector integration component to obtain the training integrated text vectors corresponding to each training text embedding vector.

[0115] Steps 4011 after step 4010 and step 4021 of step 4020 are important links for further refining the operation of the first text processing network when processing bus life - cycle file training data, aiming to better integrate information and generate a more representative bus example usage representation vector.

[0116] In step 4010, the first training text sub-vector set corresponding to each training text embedding vector has been text-precoded by the text pre-coding component to obtain the first training precoded vector. On this basis, in step 4011, the first training text sub-vector set corresponding to each training text embedding vector is position-embedded by the position embedding component to obtain the first training text position sub-vector set. The role of position embedding is to add position information to the input sub-vector set. Because when processing sequence data, the order and position of vectors often contain important semantic information. For example, in the dimension of the vehicle affairs file of a bus, the order of information such as the annual inspection time and insurance purchase time of the vehicle is very crucial for judging the usage norms and maintenance conditions of the vehicle.

[0117] To achieve position embedding, first determine the importance of the file dimension corresponding to each training text sub-vector set. The importance of the file dimension can be set according to the role of different file record dimensions in the evaluation of bus usage norms. For example, for an official vehicle of a unit, the vehicle affairs file dimension (including annual inspection, insurance, maintenance, etc.) and the violation file dimension (including violation records, penalty content, etc.) are relatively more important when evaluating whether the bus is used in a standard way. Therefore, higher priorities can be set for these two dimensions. While the importance of the vehicle file dimension (mainly the basic information of the vehicle) and the driver file dimension (basic information of the driver) may be relatively low.

[0118] After determining the importance of the file dimension, position embedding is performed on the first training text sub-vector set corresponding to each training text embedding vector through this importance. One position embedding method is to use a combination of sine and cosine functions. Assume that the dimension of the sub-vector in the first training text sub-vector set is d, the position index is pos, and the dimension index is i, then the position embedding vector . Add these position embedding vectors to the sub-vectors in the first training text sub-vector set to obtain the first training text position sub-vector set. In this way, each sub-vector contains its position information in the sequence, which helps the subsequent model to better understand the order and structure of the data.

[0119] In step 4021, the first training text position sub-vector set corresponding to each training text embedding vector is vector-integrated by the vector integration component to obtain the training integrated text vector corresponding to each training text embedding vector. The role of the vector integration component is to integrate the sub-vector set after position embedding and extract more representative features. The vector integration component can include multiple sub-components, such as an internal attention component, a first cross-layer identity mapping component and a normalization component, a multi-layer perceptron component, and a second cross-layer identity mapping component and a normalization component.

[0120] The internal attention component performs internal weight focusing on the set of first training text position sub-vectors corresponding to each training text embedding vector to obtain a first focused encoding vector. The internal attention mechanism allows the model to focus on the importance of different parts of the sub-vector set. By calculating the correlation between each sub-vector and other sub-vectors, a weight is assigned to each sub-vector, and then all sub-vectors are weighted and summed according to these weights to obtain the first focused encoding vector. For example, in the dimension of vehicle operation files, if several sub-vectors contain key annual inspection information, the internal attention mechanism will assign higher weights to these sub-vectors, thus highlighting these important information during the integration process.

[0121] Next, the first cross-layer identity mapping component and the normalization component perform a normalization operation on the first focused encoding vector and the first training pre-encoding vector corresponding to each training text embedding vector to obtain a first normalized encoding vector. The role of cross-layer identity mapping (ResNet) is to solve the problems of gradient vanishing and gradient explosion in deep networks. By directly adding the input to the output, the network can learn more complex features. The normalization operation (such as Layer Normalization) can normalize the input vector, making the training of the model more stable. Specifically, the formula for LayerNormalization is , where and are the mean and variance of the input vector, is a small constant, and are learnable parameters.

[0122] Then, the multi-layer perceptron component performs forward propagation on the first normalized encoding vector corresponding to each training text embedding vector to obtain a first forward propagation result. The multi-layer perceptron is a neural network containing multiple hidden layers, which processes the input through a series of linear transformations and non-linear activation functions (such as the ReLU function) to extract higher-level features.

[0123] Finally, the second cross-layer identity mapping component and the normalization component perform a normalization operation on the first normalized encoding vector and the first forward propagation result corresponding to each training text embedding vector to obtain the training integrated text vector corresponding to each training text embedding vector. This step again utilizes cross-layer identity mapping and normalization operations to further optimize the feature representation, so that the final training integrated text vector can more accurately reflect the usage of the bus in this file record dimension.

[0124] Through steps 4011 and 4021, when processing the first training text sub-vector set, not only the semantic information of the text is considered, but also the location information is added, and these information are effectively integrated and optimized by the vector integration component. The generated training integrated text vector can more comprehensively and accurately reflect the usage characteristics of the official vehicle, providing more powerful support for the subsequent generation of the official vehicle example usage representation vector and the evaluation of the official vehicle usage specification.

[0125] In one implementation solution, in step 4011, the first training text sub-vector set corresponding to each training text embedding vector is subjected to location embedding through the location embedding component to obtain the first training text location sub-vector set, including:

[0126] Step 40111: Determine the importance of each file dimension corresponding to the training text sub-vector set;

[0127] Step 40112: Perform location embedding on the first training text sub-vector set corresponding to each training text embedding vector according to the importance of the file dimension to obtain the first training text location sub-vector set.

[0128] In step 40111, the importance of each file dimension corresponding to the training text sub-vector set is determined. The importance of the file dimension reflects the relative importance of different file record dimensions in the evaluation of the official vehicle usage specification. Taking the official vehicles of the unit as an example, for different file record dimensions, such as the vehicle affairs file dimension, the vehicle file dimension, the driver file dimension, and the violation file dimension, they have different roles in evaluating whether the official vehicle usage is standardized. The vehicle affairs file dimension contains information such as the annual inspection, insurance, and maintenance of the vehicle. These information are directly related to the safety and legality of the vehicle and are crucial for the normal use of the official vehicle. For example, if the vehicle fails to undergo the annual inspection on time or the insurance expires, it may face legal risks and safety hazards. Therefore, the vehicle affairs file dimension has a relatively high importance in the evaluation. The vehicle file dimension mainly records the basic information of the vehicle, such as the vehicle brand, model, vehicle identification number, etc. Although these information are the basic attributes of the official vehicle, their direct role in evaluating the official vehicle usage specification is relatively small, so the importance is relatively low. The driver file dimension involves the basic information, driving qualifications, and driving records of the driver. The quality and behavior of the driver directly affect the safety and standardization of the official vehicle usage, so this dimension also has a certain importance. The violation file dimension records the violation situations of the vehicle, including the violation time, location, type, and penalty content, etc. The violation record is an important basis for evaluating whether the official vehicle usage is standardized, so the importance is relatively high.

[0129] One way to determine the importance of file dimensions is to set fixed priorities based on experience and expert knowledge. For example, set the priority of the traffic file dimension to 80%, the priority of the violation file dimension to 70%, the priority of the driver file dimension to 60%, and the priority of the vehicle file dimension to 30%. Another method is to determine the importance through data analysis. A large amount of official vehicle usage data and corresponding usage specification evaluation results can be analyzed to statistically calculate the correlation between the information of each file dimension and the evaluation results. The higher the correlation, the greater the importance of the file dimension in the evaluation. For example, the Pearson correlation coefficient is calculated to measure the linear correlation between each file dimension and the official vehicle usage specification evaluation results. The larger the absolute value of the correlation coefficient, the higher the importance of the file dimension.

[0130] In step 40112, position embedding is performed on the first training text sub-vector set corresponding to each training text embedding vector according to the file dimension importance to obtain the first training text position sub-vector set. The purpose of position embedding is to add position information to the input sub-vector set because when processing sequential data, the order and position of vectors often contain important semantic information. Position embedding can be achieved using position encoding. One position encoding method is to use a combination of sine and cosine functions. Assume that the dimension of the sub-vectors in the first training text sub-vector set is d, the position index is pos, and the dimension index is i. Then the calculation formula for the position embedding vector is:

[0131] For even dimensions (2i): ;

[0132] For odd dimensions (2i + 1): .

[0133] When performing position embedding, the position encoding will be adjusted according to the file dimension importance. For example, for file dimensions with higher importance, such as the traffic file dimension and the violation file dimension, the amplitude of the position encoding can be increased so that the sub-vectors of these dimensions are more likely to be noticed by the model in subsequent processing. Specifically, the position encoding of file dimensions with higher importance can be multiplied by a coefficient greater than 1, such as 1.2 or 1.5. For file dimensions with lower importance, such as the vehicle file dimension, the position encoding can be multiplied by a coefficient less than 1, such as 0.8 or 0.5. In this way, the position information of important file dimensions can be highlighted, and the model's ability to capture key information can be improved.

[0134] Add the adjusted positional encoding vectors to the sub-vectors in the first training text sub-vector set to obtain the first training text positional sub-vector set. In this way, each sub-vector contains the positional information in the sequence and the importance information of that file dimension. In subsequent vector integration processing, the model can better understand the order and importance of different file dimension information based on this information, so as to more accurately extract features and generate a bus usage representation vector that can comprehensively reflect the bus usage situation.

[0135] Through steps 40111 - 40112, the importance of different file dimensions is fully considered during the positional embedding process, and targeted positional information is added to the first training text sub-vector set. This not only helps the model better process sequence data, but also highlights the information of key file dimensions, improves the effect of vector integration processing, and provides a more accurate and representative input for bus usage norm evaluation.

[0136] In one implementation, the vector integration component includes an internal attention component, a first cross-layer identity mapping component and a normalization component, a multi-layer perceptron component, a second cross-layer identity mapping component and a normalization component. In step 4020, perform vector integration processing on the first training precoded vector corresponding to each training text embedding vector through the vector integration component to obtain the training integrated text vector corresponding to each training text embedding vector, including:

[0137] Step 40201: Perform internal weight focusing on the first training precoded vector corresponding to each training text embedding vector through the internal attention component to obtain the first focused encoding vector;

[0138] Step 40202: Perform a normalization operation on the first focused encoding vector and the first training precoded vector corresponding to each training text embedding vector through the first cross-layer identity mapping component and the normalization component to obtain the first normalized encoding vector;

[0139] Step 40203: Perform forward propagation on the first normalized encoding vector corresponding to each training text embedding vector through the multi-layer perceptron component to obtain the first forward propagation result;

[0140] Step 40204: Perform a normalization operation on the first normalized encoding vector and the first forward propagation result corresponding to each training text embedding vector through the second cross-layer identity mapping component and the normalization component to obtain the training integrated text vector corresponding to each training text embedding vector.

[0141] Steps 40201 - 40204 in the implementation solution of step 4020 describe the detailed process of obtaining the training integrated text vector corresponding to each training text embedding vector through vector integration processing by the vector integration component. The vector integration component includes an internal attention component, a first cross-layer identity mapping component and a normalization component, a multi-layer perceptron component, and a second cross-layer identity mapping component and a normalization component. These components work together to gradually process and optimize the input first training precoded vector, so that the final training integrated text vector can more accurately reflect the usage characteristics of the bus in each archival record dimension.

[0142] In step 40201, the internal attention component performs internal weight focusing on the first training precoded vector corresponding to each training text embedding vector to obtain the first focused coding vector. The internal attention mechanism is a technique that enables the model to automatically focus on the importance of different parts in the input sequence. In the scenario of full-life-cycle safety management of buses, the first training precoded vector contains information on multiple archival record dimensions extracted from the bus life-cycle archives, such as annual review, insurance, and maintenance information in the vehicle affairs archive dimension, basic vehicle information in the vehicle archive dimension, driver qualification information in the driver archive dimension, and violation record information in the violation archive dimension. These information may have different importance in different usage specification evaluation scenarios. For example, when evaluating whether a bus meets the safety driving standard, the annual review and maintenance information in the vehicle affairs archive dimension may be more critical; while when evaluating the compliance of bus usage, the information in the violation archive dimension is more important.

[0143] The internal attention mechanism assigns a weight to each vector by calculating the correlation between each first training precoded vector and other vectors. Specifically, the internal attention mechanism first projects the first training precoded vector into three spaces: query (Query), key (Key), and value (Value) respectively, to obtain the query matrix Q, the key matrix K, and the value matrix V. Then, by calculating the dot product of the query matrix and the key matrix, the attention score matrix S = QK T . To prevent the dot product result from being too large, S can be scaled, that is , where d_k is the dimension of the key vector. Then, the Softmax function is used to convert the attention score matrix into an attention weight matrix A = Softmax(S), so that the weight of each vector is between [0, 1] and the sum of the weights of all vectors is 1. Finally, the attention weight matrix is multiplied by the value matrix to obtain the first focused coding vector F = AV. In this way, the internal attention mechanism can automatically focus on the parts related to the current task in the input vector, highlight important information, and thus obtain a more representative first focused coding vector.

[0144] In step 40202, the first cross-layer identity mapping component and the normalization component perform a normalization operation on the first focused encoding vector and the first training pre-encoding vector corresponding to each training text embedding vector to obtain the first normalized encoding vector. Cross-layer identity mapping (ResNet) is a technique used to address the problems of vanishing gradients and exploding gradients in deep neural networks. During the vector integration process, as the number of network layers increases, the gradients may become very small or very large during backpropagation, making it difficult to train the model. Cross-layer identity mapping alleviates the problems of vanishing gradients and exploding gradients by directly adding the input to the output, enabling the gradients to be directly propagated from the output layer to the input layer. Specifically, the first cross-layer identity mapping component adds the first training pre-encoding vector X and the first focused encoding vector F to obtain R = X + F.

[0145] The normalization component then performs a normalization process on the added result to improve the training stability of the model. The normalization method is, for example, Layer Normalization, which has been introduced above and will not be elaborated here. Through Layer Normalization, each dimension of the input vector is normalized to the same scale, enabling the model to maintain a stable training effect under different inputs. After being processed by the first cross-layer identity mapping component and the normalization component, the first normalized encoding vector N1 = LN(R) is obtained.

[0146] In step 40203, the multi-layer perceptron component performs forward propagation on the first normalized encoding vector corresponding to each training text embedding vector to obtain the first forward propagation result. A multi-layer perceptron (MLP) is a neural network with multiple hidden layers that can learn complex non-linear relationships between input vectors. In the full-life-cycle safety management of buses, the first normalized encoding vector contains information that has undergone internal attention focusing and normalization, but this information may require further feature extraction and transformation to be better used for bus usage specification evaluation. The multi-layer perceptron component processes the first normalized encoding vector through a series of linear transformations and non-linear activation functions.

[0147] Assume that the multi-layer perceptron has L hidden layers, and the input of the l-th layer is h l-1 (where h 0 = N1), then the output of the l-th layer is , where W l is the weight matrix of the l-th layer, and b lis the bias vector of the l-th layer, and f is a non-linear activation function, such as the ReLU function (f(x) = max(0, x)). The ReLU function can introduce non-linearity, enabling the multi-layer perceptron to learn more complex feature representations. After the forward propagation of the multi-layer perceptron, the first forward propagation result P is finally obtained.

[0148] In step 40204, the second cross-layer identity mapping component and the normalization component perform a normalization operation on the first normalized encoding vector and the first forward propagation result corresponding to each training text embedding vector to obtain the training integrated text vector corresponding to each training text embedding vector. Similar to step 40202, the second cross-layer identity mapping component adds the first normalized encoding vector N1 and the first forward propagation result P to get R' = N1 + P. This step again utilizes cross-layer identity mapping to directly transfer the information of the previous layer to the current layer, which helps to alleviate the problems of gradient vanishing and gradient explosion and retain more original information.

[0149] Then, the normalization component performs Layer Normalization on the added result. After the processing of the second cross-layer identity mapping component and the normalization component, the training integrated text vector T = LN(R') corresponding to each training text embedding vector is obtained.

[0150] Through steps 40201 - 40204, a series of processing and optimization are performed on the first training precoding vector using the vector integration component. The internal attention component focuses on important information, the cross-layer identity mapping component and the normalization component improve the training stability of the model, and the multi-layer perceptron component learns the complex non-linear relationship between the input vectors. The finally obtained training integrated text vector can more accurately reflect the usage characteristics of the bus under each archival record dimension, providing stronger support for the subsequent generation of the bus sample usage representation vector and the evaluation of bus usage specifications.

[0151] As another implementation solution for training the first text processing network and the second text processing network, the solution includes:

[0152] Step 1000: Obtain archival record training texts, the supervision marks of the usage specification evaluation results corresponding to the record entities in the archival record training texts, and the relevance supervision marks of the archival record training texts; the relevance supervision marks are used to represent the task execution information of the relevance task of the archival record training texts;

[0153] Step 2000: Load the archival record training texts into the first text processing network for text vector embedding to obtain the training text vectors corresponding to the archival record training texts;

[0154] Step 3000: Load the training text vectors into the second text processing network, perform a fully connected classification mapping on the archival record training text, and obtain the usage specification evaluation results corresponding to the record entities in the archival record training text;

[0155] Step 4000: Load the training text vectors into the relevance task network, perform relevance task processing on the archival record training text, and obtain the relevance task processing results corresponding to the archival record training text;

[0156] Step 5000: Determine the training error value based on the usage specification evaluation results and the usage specification evaluation result supervision labels, as well as the relevance task processing results and the relevance supervision labels;

[0157] Step 6000: Based on the training error value, perform convergence training on the network parameters of the first text processing network and the second text processing network.

[0158] In step 1000, obtain the archival record training text, the usage specification evaluation result supervision labels corresponding to the record entities in the archival record training text, and the relevance supervision labels of the archival record training text. The archival record training text is a text collection on the full life cycle management of official vehicles, covering information in multiple dimensions such as vehicle operation files, vehicle files, driver files, and violation files. The usage specification evaluation result supervision labels clarify the results of whether the use of official vehicles corresponding to each archival record training text is standardized, such as "standardized", "basically standardized", "non-standardized", etc. The relevance supervision labels are used to represent the task execution information of the relevance tasks of the archival record training text. For example, in some scenarios, it is necessary to judge the relevance between different information in the archival record training text, such as whether there is an association between the annual review information in the vehicle operation file and the speeding violation information in the violation file. The relevance supervision labels will record this association information. Taking the official vehicles of a unit as an example, the archival record training text may include detailed records of a vehicle's certain maintenance, driver's violation ticket information, etc. The usage specification evaluation result supervision labels will indicate whether the use of the vehicle during this period is in line with the specifications, and the relevance supervision labels may record whether there is a potential connection between the maintenance time and the violation time.

[0159] In step 2000, the archival record training text is loaded into the first text processing network for text vector embedding to obtain the training text vector corresponding to the archival record training text. The first text processing network can adopt pre-trained language models such as BERT, XLNet, etc. These models can learn the semantic information and context relationship of the text, and convert the input archival record training text into a low-dimensional vector representation. Taking BERT as an example, it encodes the input text through a multi-layer Transformer architecture. Suppose the archival record training text is a document containing multiple sentences. BERT will convert the words in each sentence into word vectors, then calculate the correlation between words through the self-attention mechanism, and finally output the training text vector of the entire document. Specifically, the input of BERT is obtained by adding word embedding, position embedding, and segment embedding. After being processed by multiple Transformer blocks, the context representation of each word is obtained, and the output corresponding to the [CLS] token is taken as the training text vector of the entire document.

[0160] In step 3000, the training text vector is loaded into the second text processing network to perform a fully connected classification mapping on the archival record training text, and obtain the usage specification evaluation result corresponding to the record entity in the archival record training text. The second text processing network can be a multi-layer perceptron (MLP), which consists of an input layer, a hidden layer, and an output layer. The input layer receives the training text vector output by the first text processing network. The hidden layer processes the input through a series of linear transformations and non-linear activation functions (such as the ReLU function) to extract higher-level features. The number of neurons in the output layer is the same as the number of categories of the usage specification evaluation result. The output is converted into a probability distribution through the Softmax function. Suppose the usage specification evaluation result is divided into three cases: "standard", "basic standard", and "non-standard". The output layer has 3 neurons. After being processed by the Softmax function, the obtained probability distribution P = [p1, p2, p3] respectively represents the probabilities that the bus usage situation corresponding to the archival record training text is "standard", "basic standard", and "non-standard". The category with the highest probability is selected as the final usage specification evaluation result.

[0161] In step 4000, the training text vector is loaded into the correlation task network to perform correlation task processing on the archival record training text, and obtain the correlation task processing result corresponding to the archival record training text. The correlation task network can be a specially designed neural network for processing the correlation between different information in the archival record training text. For example, it can judge whether they are relevant by calculating the similarity between the vectors of different dimensional information. Suppose the vector corresponding to the information in the vehicle affairs file dimension is , and the vector corresponding to the information in the violation file dimension is , the relevance task network can use cosine similarity to calculate the similarity between them , and determine whether the information in these two dimensions is relevant according to the value of the similarity. The relevance task network processes the training text vectors and outputs the relevance task processing results, such as relevant or irrelevant.

[0162] In step 5000, based on the usage specification evaluation result and the usage specification evaluation result supervision label, as well as the relevance task processing result and the relevance supervision label, the training error value is determined. For the usage specification evaluation task, the cross-entropy loss function L1 is used to measure the difference between the predicted usage specification evaluation result and the actual usage specification evaluation result supervision label. The formula of the cross-entropy loss function refers to the description in step 50 above and will not be elaborated here. For the relevance task, the mean squared error loss function can be used to measure the difference between the relevance task processing result and the relevance supervision label. The formula of the mean squared error loss function is , where r j is the j-th element of the relevance supervision label, is the j-th element of the relevance task processing result, and n is the number of samples. The final training error value L = L1 + L2, which combines the errors of the usage specification evaluation task and the relevance task.

[0163] In step 6000, based on the training error value, the network parameters of the first text processing network and the second text processing network are trained for convergence. The backpropagation algorithm is used to calculate the gradients of the training error value with respect to the network parameters (such as the weight matrix and the bias vector), and then the parameters are updated according to the gradients. Feasible optimization algorithms include Stochastic Gradient Descent (SGD), Adagrad, Adadelta, Adam, etc. Through multiple iterations of training, the network parameters are continuously adjusted to gradually reduce the training error value until the convergence condition is reached (such as the error value is less than a certain threshold or the number of iterations reaches the upper limit). At this time, the performance of the network reaches the optimal, and it can accurately evaluate the usage specifications of the bus and process the relevance tasks in the archival record training text. Through this training scheme, the usage specification evaluation task and the relevance task are combined, enabling the first text processing network and the second text processing network to learn richer information during the training process. The relevance task can help the network better understand the relationships between different pieces of information in the archival record training text, thereby improving the accuracy of the usage specification evaluation. At the same time, through joint training, the network can share information and promote each other between the two tasks, making the final model have better performance in the full life cycle safety management of the bus.

[0164] In the above training scheme, a multi-task learning mechanism is adopted. When training the network, the archival record training text for training the network and its corresponding supervision label of the usage specification evaluation result are obtained, and the relevance supervision label of the archival record training text is also obtained. Then, after extracting the text vector of the archival record training text, the obtained training text vector is loaded into the second text processing network for evaluation and classification, and is loaded into the relevance task network for performing association tasks. Then, after fusing the training error values obtained from the two tasks, the network is trained for convergence. In this way, a training mechanism of multi-task collaboration can be adopted to improve the semantic understanding effect of the network, and further improve the accuracy of the specification evaluation result.

[0165] An embodiment of the present application provides a computer system, including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, some or all of the steps in the above method are implemented.

[0166] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, some or all of the steps in the above method are implemented. The computer-readable storage medium can be transient or non-transient.

[0167] An embodiment of the present application provides a computer program, including computer-readable code. When the computer-readable code runs in a computer device, the processor in the computer device executes to implement some or all of the steps in the above method.

[0168] An embodiment of the present application provides a computer program product. The computer program product includes a non-transient computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. The computer program product can be specifically implemented in a way of hardware, software or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium. In other embodiments, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0169] It should be pointed out here that the descriptions of the above embodiments tend to emphasize the differences between the embodiments, and their similarities or similarities can be referred to each other. The descriptions of the above embodiments of the device, storage medium, computer program and computer program product are similar to the descriptions of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the embodiments of the device, storage medium, computer program and computer program product of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.

[0170] Figure 2 The following is a schematic diagram of the hardware entity of a computer system provided by an embodiment of the present application. As Figure 2 shown, the hardware entity of the computer system 1000 includes: a processor 1001 and a memory 1002. Among them, the memory 1002 stores a computer program that can run on the processor 1001. When the processor 1001 executes the program, it implements the steps in the method of any of the above embodiments.

[0171] The memory 1002 stores a computer program that can run on the processor. The memory 1002 is configured to store instructions and applications executable by the processor 1001, and can also cache data to be processed or already processed by the processor 1001 and each module in the computer system 1000 (for example, image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM).

[0172] When the processor 1001 executes the program, it implements the steps of the bus full-life-cycle safety management method combined with big data in any of the above items. The processor 1001 controls the overall operation of the computer system 1000.

[0173] An embodiment of the present application provides a computer storage medium. The computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the bus full-life-cycle safety management method combined with big data in any of the above embodiments.

[0174] It should be noted here that the descriptions of the above storage medium and device embodiments are similar to those of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the storage medium and device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding. The above-mentioned processor may be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that other electronic devices for implementing the functions of the above-mentioned processor are also possible, and the embodiments of the present application do not make specific limitations.

[0175] The above computer storage medium / memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM), etc.; it may also be various terminals including one or any combination of the above memories, such as a mobile phone, a computer, a tablet device, a personal digital assistant, etc.

[0176] It should be understood that in various embodiments of the present application, the magnitude of the serial numbers of the above steps / processes does not mean the order of execution. The order of execution of each step / process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments. It should be noted that in this text, the terms "include", "comprise" or any other variation thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.

[0177] In addition, in each embodiment of the present application, each functional unit can be all integrated in a processing unit, or each unit can be separately taken as a unit alone, or two or more units can be integrated in a unit; the above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0178] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memory (ROM), magnetic disks or optical disks and other various media that can store program codes.

[0179] Alternatively, if the above integrated unit of the present application is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present application essentially or the part that contributes to the related technology can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. And the foregoing storage medium includes: removable storage devices, ROM, magnetic disks or optical disks and other various media that can store program codes.

[0180] The above is only the implementation mode of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, and all should be covered within the protection scope of the present application.

Claims

1. A method for the whole life cycle safety management of buses combined with big data, characterized in that, The method includes: Obtaining a historical record text of a bus life cycle file, where the bus life cycle file is a record set for the full life cycle management of a target bus; Extracting archival text segments of multiple archival record dimensions from the historical record text, and mapping the extracted multiple archival text segments into multiple text embedding vectors; wherein, the multiple archival record dimensions include a vehicle affairs file dimension, a vehicle file dimension, a driver file dimension, and a violation file dimension; Taking the first vector coverage range as the window size and the second vector coverage range as the window stride, performing a sliding window process on each of the text embedding vectors to obtain a first set of text sub-vectors corresponding to each of the text embedding vectors; Performing vector integration processing on each of the first sets of text sub-vectors, and combining the multiple integrated text vectors obtained by the integration to obtain a bus usage characterization vector of the bus life cycle file; Determining an evaluation result of the usage specification of the bus life cycle file through the bus usage characterization vector; Among them, the performing vector integration processing on each of the first sets of text sub-vectors and combining the multiple integrated text vectors obtained by the integration to obtain a bus usage characterization vector of the bus life cycle file includes: determining a first text processing network corresponding to each of the first sets of text sub-vectors; loading the first set of text sub-vectors corresponding to each of the text embedding vectors into the corresponding first text processing network for vector integration processing to obtain an integrated text vector corresponding to each of the text embedding vectors; determining an evaluation contribution coefficient of each of the integrated text vectors; performing weighted adjustment on each of the integrated text vectors through the evaluation contribution coefficient to obtain multiple adjusted text vectors; performing vector combination on the multiple adjusted text vectors to obtain a bus usage characterization vector of the bus life cycle file; wherein, the process of determining the evaluation contribution coefficient includes: obtaining a vehicle empowerment level; determining the evaluation contribution coefficient of each of the integrated text vectors according to the correlation information between each of the archival record dimensions and the vehicle empowerment level; After extracting archival text segments of multiple archival record dimensions from the historical archival record text and mapping the multiple obtained archival text segments into multiple text embedding vectors, the method further includes: using the third vector coverage range as the window size and the fourth vector coverage range as the window stride, performing a sliding window process on each of the text embedding vectors to obtain a corresponding second set of text sub-vectors for each of the text embedding vectors; at this time, the vector integration process for each of the first sets of text sub-vectors and the combination of the multiple integrated text vectors obtained through integration to obtain the bus usage characterization vector of the bus life cycle file includes: performing a vector integration process on each of the first sets of text sub-vectors and combining the multiple first integrated sub-text vectors obtained through integration to obtain a first bus usage sub-characterization vector, performing a vector integration process on each of the second sets of text sub-vectors, and combining the multiple second integrated sub-text vectors obtained through integration to obtain a second bus usage sub-characterization vector; combining the first bus usage sub-characterization vector and the second bus usage sub-characterization vector to obtain the bus usage characterization vector of the bus life cycle file.

2. The method according to claim 1, characterized in that, Determining the usage specification evaluation result of the bus life cycle file through the bus usage characterization vector includes: Loading the bus usage characterization vector into a second text processing network for fully connected mapping to obtain a first evaluation probability distribution; Determining the usage specification evaluation result of the bus life cycle file according to the first evaluation probability distribution.

3. The method according to claim 2, wherein The multiple first text processing networks and the second text processing network are obtained through parallel training using bus life cycle file samples. The multiple first text processing networks and the second text processing network are trained through the following process: Obtaining archival record training text, where the archival record training text includes full-cycle archival record training text of multiple bus training examples and supervision labels for the usage specification evaluation results of each bus training example; Extracting archival training text segments of multiple archival record dimensions from the full-cycle archival record training text and mapping the multiple obtained archival training text segments into multiple training text embedding vectors; Using the first vector coverage range as the window size and the second vector coverage range as the window stride, performing a sliding window process on each of the training text embedding vectors to obtain a corresponding first set of training text sub-vectors for each of the training text embedding vectors; Loading the first set of training text sub-vectors corresponding to each of the training text embedding vectors into the corresponding first text processing network for vector integration processing, and combining the training integrated text vectors corresponding to each of the training text embedding vectors obtained through integration to obtain a bus example usage characterization vector; Loading the bus example usage characterization vector into a second text processing network for fully connected mapping to obtain a second evaluation probability distribution, and determining a mapping error value through the second evaluation probability distribution and the corresponding supervision label for the usage specification evaluation result; Converge and train the network parameters of multiple said first text processing networks and said second text processing networks through the mapping error value.

4. The method according to claim 3, wherein After extracting archival training text segments of multiple archival record dimensions from the full-cycle archival record training text and mapping the extracted multiple said archival training text segments into multiple training text embedding vectors, the method further includes: Taking the third vector coverage range as the window size and the fourth vector coverage range as the window stride, perform sliding window processing on each said training text embedding vector to obtain a second set of training text sub-vectors corresponding to each said training text embedding vector; Loading the first set of training text sub-vectors corresponding to each said training text embedding vector into the corresponding said first text processing network for vector integration processing, and combining the training integrated text vectors corresponding to each said training text embedding vector obtained by integration to obtain a bus example usage representation vector, including: Loading the first set of training text sub-vectors corresponding to each said training text embedding vector into the corresponding said first text processing network for vector integration processing, and combining the first training integrated sub-text vectors corresponding to each said training text embedding vector obtained by integration to obtain a first training bus usage sub-representation vector; Loading the second set of training text sub-vectors corresponding to each said training text embedding vector into the corresponding said first text processing network for vector integration processing, and combining the second training integrated sub-text vectors corresponding to each said training text embedding vector obtained by integration to obtain a second training bus usage sub-representation vector; Combining the first training bus usage sub-representation vector and the second training bus usage sub-representation vector to obtain the bus example usage representation vector of the bus training example; Each said first text processing network includes a text pre-coding component and a vector integration component. Loading the first set of training text sub-vectors corresponding to each said training text embedding vector into the corresponding said first text processing network for vector integration processing, and combining the training integrated text vectors corresponding to each said training text embedding vector obtained by integration to obtain a bus example usage representation vector, including: Performing text pre-coding on the first set of training text sub-vectors corresponding to each said training text embedding vector through the text pre-coding component to obtain a first training pre-coded vector; Performing vector integration processing on the first training pre-coded vector corresponding to each said training text embedding vector through the vector integration component to obtain a training integrated text vector corresponding to each said training text embedding vector; Combining the training integrated text vectors corresponding to each said training text embedding vector to obtain a bus example usage representation vector.

5. The method according to claim 4, characterized in that, Each said first text processing network further includes a position embedding component. After performing text pre-coding on the first set of training text sub-vectors corresponding to each said training text embedding vector through the text pre-coding component to obtain a first training pre-coded vector, the method further includes: Performing positional embedding on the first training text sub-vector set corresponding to each of the training text embedding vectors through the position embedding component to obtain a first training text position sub-vector set; The performing vector integration processing on the first training precoding vector corresponding to each of the training text embedding vectors through the vector integration component to obtain a training integrated text vector corresponding to each of the training text embedding vectors includes: Performing vector integration processing on the first training text position sub-vector set corresponding to each of the training text embedding vectors through the vector integration component to obtain a training integrated text vector corresponding to each of the training text embedding vectors.

6. The method according to claim 5, wherein The performing positional embedding on the first training text sub-vector set corresponding to each of the training text embedding vectors through the position embedding component to obtain a first training text position sub-vector set includes: Determining the importance of each file dimension corresponding to the training text sub-vector set; Performing positional embedding on the first training text sub-vector set corresponding to each of the training text embedding vectors through the importance of the file dimension to obtain a first training text position sub-vector set.

7. The method according to claim 4, characterized in that, The vector integration component includes an internal attention component, a first cross-layer identity mapping component and a normalization component, a multi-layer perceptron component, a second cross-layer identity mapping component and a normalization component. The performing vector integration processing on the first training precoding vector corresponding to each of the training text embedding vectors through the vector integration component to obtain a training integrated text vector corresponding to each of the training text embedding vectors includes: Performing internal weight focusing on the first training precoding vector corresponding to each of the training text embedding vectors through the internal attention component to obtain a first focused coding vector; Performing a normalization operation on the first focused coding vector and the first training precoding vector corresponding to each of the training text embedding vectors through the first cross-layer identity mapping component and the normalization component to obtain a first normalized coding vector; Performing forward propagation on the first normalized coding vector corresponding to each of the training text embedding vectors through the multi-layer perceptron component to obtain a first forward propagation result; Performing a normalization operation on the first normalized coding vector and the first forward propagation result corresponding to each of the training text embedding vectors through the second cross-layer identity mapping component and the normalization component to obtain a training integrated text vector corresponding to each of the training text embedding vectors.

8. A computer system, comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Malicious comment detection method based on Bert and Bi-LSTM

    CN116881449A