Big data-combined bus full-life-cycle safety management method and system

By combining big data methods, text information from multiple archive record dimensions in the bus life cycle archive is extracted and integrated, and bus usage representation vectors are generated, which solves the problem that existing evaluation methods are difficult to accurately identify the irregular use of buses, and improves the accuracy and reliability of the evaluation.

CN119941480AActive Publication Date: 2025-05-06GUIYANG JINYANG CONSTR DATA SERVICE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510421869.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The existing standardized use evaluation methods for buses are difficult to comprehensively and accurately capture the information in the archive record dimensions, resulting in the difficulty of effectively identifying the single-dimensional irregular vehicle use behavior, affecting the accuracy of the evaluation.

Method used

The full life cycle safety management method of buses combined with big data is adopted. By obtaining the historical archive record text of bus life cycle archives, text fragments of multiple archive record dimensions are extracted, mapped into text embedding vectors, and moving ruler processing and vector integration are carried out to generate bus usage characterization vectors to evaluate usage specifications.

Benefits of technology

It improves the accuracy and reliability of the evaluation of standardized use of buses, enhances the ability to identify single-dimensional irregular vehicle use behaviors, and ensures that the use of buses meets the requirements of the specifications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941480A_ABST
    Figure CN119941480A_ABST
Patent Text Reader

Abstract

The invention provides a public vehicle full life cycle safety management method and system in combination with big data. The method comprises the steps of obtaining a historical archive record text of a bus life cycle archive; extracting archive text fragments of a plurality of archive record dimensions from the historical archive record text, and mapping the plurality of extracted archive text fragments into a plurality of text embedding vectors; taking the first vector coverage range as a measurement size, taking the second vector coverage range as a measurement stride, and performing mobile measurement processing on each text embedding vector to obtain a first text sub-vector set corresponding to each text embedding vector; performing vector integration processing on each first text sub-vector set, and combining a plurality of integrated text vectors obtained by integration to obtain a bus use representation vector of the bus life cycle file; and determining a use specification evaluation result of the bus life cycle file through the bus use representation vector. According to the invention, the accuracy of bus use specification evaluation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of text processing technology, and in particular to a method and system for bus life cycle safety management combined with big data. Background Art

[0002] In the daily operation of enterprises and institutions, public buses are important means of transportation for performing official duties, and their standardized use is of great significance. The standardized use of public buses faces many problems and challenges. The phenomenon of using public buses for private purposes occurs from time to time. Some personnel violate regulations and use public buses for private affairs, resulting in serious waste of public resources; the problem of over-standard configuration is also prominent. Some units fail to configure public buses according to the prescribed standards, resulting in irrational use of resources; there are also deficiencies in vehicle maintenance and management. A considerable number of public buses are not maintained and serviced on time, resulting in poor vehicle condition. At the same time, the cost of vehicle repair and maintenance is high, and the management method needs to be further optimized. The existing evaluation methods for the standardized use of public buses have certain limitations. Traditional evaluation methods are often difficult to fully and accurately capture the information of the archive record dimension. The normative evaluation features of a single archive dimension that is relatively concentrated in a short period of time are easily ignored in the global text embedding vector, making it difficult to effectively identify the irregular use of vehicles in a single dimension, which in turn affects the accuracy of the evaluation of the standardized use of public buses. Therefore, a more effective technical solution is needed to improve the accuracy and reliability of the evaluation of the standardized use of public buses, so as to better strengthen the management of public buses, ensure that the use of public buses complies with the standardized requirements, and give full play to the role of public buses in official activities. Summary of the invention

[0003] In view of this, the embodiment of the present application at least provides a bus life cycle safety management method and system combined with big data. The technical solution of the present application is implemented as follows: On the one hand, an embodiment of the present application provides a method for managing the safety of a bus throughout its life cycle in combination with big data, the method comprising: obtaining a historical archive record text of a bus life cycle archive, the bus life cycle archive being a record set for the full life cycle management of a target bus; extracting archive text fragments of multiple archive record dimensions from the historical archive record text, and mapping the extracted multiple archive text fragments into multiple text embedding vectors; wherein the multiple archive record dimensions include a vehicle service archive dimension, a vehicle archive dimension, a driver archive dimension, and a violation archive dimension; using the first vector coverage range as a scaling size, and using the second vector coverage range as a scaling stride, performing a moving scaling process on each of the text embedding vectors to obtain a first text sub-vector set corresponding to each of the text embedding vectors; performing vector integration processing on each of the first text sub-vector sets, and combining the multiple integrated text vectors obtained by integration to obtain a bus usage representation vector of the bus life cycle archive; determining a usage specification evaluation result of the bus life cycle archive through the bus usage representation vector.

[0004] On the other hand, the present application provides a computer system, including a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor implements the steps in the above method when executing the program.

[0005] The present application provides a bus life cycle safety management method combined with big data, which obtains historical archive record texts of bus life cycle archives; extracts archive text fragments of multiple archive record dimensions from the historical archive record texts, and maps the extracted multiple archive text fragments into multiple text embedding vectors; uses the first vector coverage range as the scaling size and the second vector coverage range as the scaling step, performs mobile scaling processing on each text embedding vector, and obtains a first text sub-vector set corresponding to each text embedding vector; performs vector integration processing on each first text sub-vector set, and combines the multiple integrated text vectors obtained by integration to obtain a bus use representation vector of the bus life cycle archive; and determines the use specification evaluation result of the bus life cycle archive through the bus use representation vector.

[0006] Through the above scheme, this application adopts the technology of mobile scale processing to extract multiple scale features (that is, the fusion result of text vectors of several attribute dimensions) from the text embedding vector of each archive record dimension, so that a vector combination with richer information can be obtained for each archive record dimension. Then, it can prevent the normative evaluation features of a single archive dimension that is relatively concentrated in a short period of time from being ignored in the global text embedding vector, thereby increasing the recognition ability of single-dimensional irregular use of vehicles and improving the accuracy of the evaluation of the standardized use of public vehicles.

[0007] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 A schematic diagram of the implementation process of a method for managing the safety of a bus throughout its life cycle in combination with big data provided in an embodiment of the present application.

[0009] Figure 2 A hardware entity diagram of a computer system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0010] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions of the present application are further elaborated in detail below in conjunction with the drawings and embodiments. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0011] The embodiment of the present application provides a bus life cycle safety management method combined with big data, which can be executed by a processor of a computer system. The computer system can refer to a server, a laptop, a tablet computer, a desktop computer, a mobile device (such as a mobile phone, a portable video player, a personal digital assistant), and other devices with data processing capabilities.

[0012] Figure 1 A schematic diagram of the implementation process of a bus life cycle safety management method combined with big data provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the method includes: Step 100: Obtain the historical archive record text of the bus life cycle archive, where the bus life cycle archive is a collection of records for the full life cycle management of the target bus.

[0013] Official vehicles are official vehicles, such as vehicles used by enterprises and institutions to perform official activities. The whole life cycle management covers the entire process of official vehicles from purchase, use, maintenance to scrapping, including vehicle archive dimension information, such as annual inspection, insurance, maintenance, etc.; vehicle archive dimension information, such as vehicle basic information related information; driver archive dimension information, such as driver basic information related information; and violation archive dimension information, such as violation records, violators, punishment content, etc. The historical archive record text is a collection of these whole life cycle management information recorded in text form.

[0014] For example, for a company's official vehicle, its life cycle archive may include: vehicle archive information such as the vehicle purchase invoice, vehicle model and configuration information; annual inspection report, insurance contract and other vehicle archive information; driver's license, work history and other driver archive information; and violation archive information such as violation tickets and handling results generated during the use of the vehicle. This information is stored in the form of documents, reports, tables, etc., and needs to be collected and organized into historical archive record texts for subsequent processing.

[0015] In order to obtain the text of these historical archive records, the following methods can be used. If these archive records are stored in paper documents, optical character recognition (OCR) technology can be used to convert the text in the paper documents into electronic text. OCR technology recognizes and analyzes the characters in the image and converts them into a text format that can be processed by a computer. For example, for a paper annual review report, a scanner can be used to convert it into an image file, and then the OCR software can be used to recognize the text in the image, and finally the annual review report in electronic text form can be obtained.

[0016] If the archival records are stored in electronic documents, these documents can be read directly. For structured electronic documents, such as tables in a database, SQL statements can be used to query and extract data. For unstructured electronic documents, such as Word documents and PDF documents, the corresponding document parsing library can be used for text extraction. For example, for a Word document, Python's python-docx library can be used to read the document content and convert it into text form.

[0017] Step 200: extracting archive text segments of multiple archive record dimensions from the historical archive record text, and mapping the extracted multiple archive text segments into multiple text embedding vectors; wherein the multiple archive record dimensions include a vehicle service archive dimension, a vehicle archive dimension, a driver archive dimension, and a violation archive dimension.

[0018] Multiple file record dimensions include the vehicle service file dimension, which covers annual inspection, insurance, maintenance and other information; the vehicle file dimension, which involves information related to the basic vehicle information; the driver file dimension, which contains information related to the basic driver information; and the violation file dimension, such as violation records, violators, and penalty content.

[0019] For example, for the official vehicle of the unit mentioned above, its historical archive record text may be a comprehensive document containing various information. It is necessary to extract archive text fragments of different archive record dimensions. For the vehicle archive dimension, text fragments such as "the annual vehicle inspection was completed on October 15, 2024, and all indicators were qualified", "the vehicle insurance period is from January 1, 2024 to December 31, 2024, and the insurance amount is 500,000 yuan", "on July 20, 2024, the vehicle was repaired due to engine failure, and the repair cost was 3,000 yuan" will be extracted. For the vehicle archive dimension, texts such as "the vehicle brand is Volkswagen, the model is Passat, and the frame number is LSVCC2A46FN012345" will be extracted. For the driver archive dimension, content such as "the driver's name is Zhang San, the driver's license number is 1234567890, and the permitted vehicle type is C1" will be extracted. For the traffic violation file dimension, text fragments such as "On August 5, 2024, driver Li Si exceeded the speed limit on Zhongshan Road and was fined 200 yuan and deducted 3 points" will be extracted.

[0020] In order to extract these archival text fragments, a rule-based approach can be used. By defining a series of rules, such as keyword matching, regular expressions, etc., text fragments of different archival record dimensions can be located and extracted. For example, for the annual inspection information of the vehicle service archive dimension, the rule can be defined to match sentences containing keywords such as "annual inspection" and "annual inspection". For the vehicle model information of the vehicle archive dimension, regular expressions can be used to match vehicle models in a specific format. Another method is a machine learning-based method, such as a named entity recognition (NER) model. A pre-trained NER model can be used to process historical archival record texts and identify entity information of different archival record dimensions. For example, using a pre-trained model such as BERT, through fine-tuning training, it can be enabled to recognize entities of different dimensions such as vehicle service archives, vehicle archives, driver archives, and violation archives.

[0021] After extracting archival text fragments of multiple archival record dimensions, these archival text fragments are mapped into multiple text embedding vectors. The text embedding vector converts the text into a representation in the vector space for subsequent calculations and analysis. Text embedding methods include word vector models, such as Word2Vec, GloVe, etc., and language models based on deep learning, such as BERT, XLNet, etc. Taking Word2Vec as an example, each word is mapped to a vector space of fixed dimension by training a neural network model. For an archival text fragment, each word in it can be converted into a corresponding word vector, and then these word vectors can be combined into a text embedding vector by averaging, summing, etc. Assuming that an archival text fragment is "Vehicle Annual Inspection Passed", first convert the three words "vehicle", "annual inspection" and "passed" into corresponding word vectors respectively. , and then the text embedding vector of the text fragment is obtained by averaging .

[0022] For language models based on deep learning, such as BERT, the entire text fragment can be directly encoded to obtain a fixed-dimensional text embedding vector. The archival text fragment is input into the pre-trained BERT model, and the model will output the text embedding vector corresponding to the text fragment. By mapping archival text fragments of multiple archival record dimensions into multiple text embedding vectors, the text information is converted into a computer-processable vector form, which provides a basis for subsequent mobile scale processing, vector integration processing and other operations, and helps to more accurately analyze the use of buses and conduct standardized evaluations.

[0023] Step 300: Using the first vector coverage as the scaling size and the second vector coverage as the scaling step, perform moving scaling processing on each text embedding vector to obtain a first text sub-vector set corresponding to each text embedding vector.

[0024] Moving scaling is a method for local feature extraction on sequence data. By setting the scaling size and scaling stride, a window is slid on the text embedding vector, and each time the vector in the window is intercepted as a sub-vector, thus obtaining a series of sub-vector sets.

[0025] Taking the official vehicles mentioned above as an example, we have mapped the archive text fragments of different archive record dimensions into multiple text embedding vectors. Assume that a text embedding vector of the vehicle archive dimension is , each of the is a numerical vector, representing a point of the text fragment in the vector space.

[0026] The first vector coverage is the scale size, which determines the length of the vector each time it is intercepted. Assume that the first vector coverage is set to 3, and the second vector coverage is set to 2. Starting from the starting position of the text embedding vector, the vector is intercepted with a scale size of 3 to obtain the first sub-vector Then move the window with a stride of 2 and intercept the second sub-vector , continue to move the window to get , and so on, until no more sub-vectors that meet the scale can be intercepted. Finally, for this text embedding vector of the vehicle service file dimension, the first text sub-vector set obtained is .

[0027] In actual operation, the moving scale processing can be implemented by loop traversal. For each text embedding vector, starting from the first element of the vector, the scale operation is performed according to the scale size and scale stride. Suppose the text embedding vector is , the size is k, the stride is s, then the starting position i of the subvector starts from 0 and increases by s each time until i+k>n. The jth subvector ,in .

[0028] This method of moving scale processing can extract multiple scale features from the text embedding vector of each archive record dimension by setting the appropriate scale size and scale stride. These scale features are the fusion results of text vectors of several attribute dimensions, which can make the vector combination information of each archive record dimension richer. For example, in the vehicle service archive dimension, a relatively concentrated vehicle maintenance information in a short period of time may be ignored in the global text embedding vector, but through moving scale processing, this part of information can be extracted separately to form a valuable sub-vector, thereby increasing the recognition ability of irregular vehicle use in a single dimension and improving the accuracy of the evaluation of the standardized use of public vehicles.

[0029] Step 400: Perform vector integration processing on each first text sub-vector set, and combine multiple integrated text vectors obtained by integration to obtain a bus usage representation vector of the bus life cycle file.

[0030] In step 400, vector integration processing is performed on each first text sub-vector set, and multiple integrated text vectors obtained by integration are combined to obtain a bus usage representation vector of the bus life cycle archive. Vector integration processing is to merge and transform the first text sub-vector set corresponding to each archive record dimension to extract more representative features. The bus usage representation vector is a vector representation that comprehensively reflects the usage of the bus throughout its life cycle.

[0031] Taking the official vehicles mentioned above as an example, the text embedding vectors of different archive record dimensions have been processed by moving scale, and the first text sub-vector set corresponding to each dimension has been obtained. For example, the first text sub-vector set of the vehicle service archive dimension is , the first text sub-vector set of the vehicle file dimension is , the first text sub-vector set of the driver profile dimension is , the first text sub-vector set of the violation file dimension is .

[0032] There are many ways to perform vector integration. One method is to use a neural network model, such as a long short-term memory network (LSTM) or a gated recurrent unit (GRU). These models can process sequence data, model the first text sub-vector set, and capture the time series information therein. Taking LSTM as an example, each sub-vector in the first text sub-vector set is input into the LSTM model in turn. The model updates the hidden state based on the current input and the hidden state at the previous moment, and finally outputs an integrated vector. Suppose the first text sub-vector set is , the hidden state update formula of the LSTM model is , where h t is the hidden state at time t, V t is the input subvector at time t, h t-1 is the hidden state of the previous moment. After processing for m moments, the final hidden state h m is the integrated vector.

[0033] Another method is to use the attention mechanism. The attention mechanism allows us to pay attention to the importance of different sub-vectors when integrating vectors. We calculate the attention weight of each sub-vector and then perform a weighted summation of all sub-vectors based on these weights. Let the first text sub-vector set be , the attention weight is , then the integrated vector ,in .

[0034] After completing the vector integration process of each archive record dimension, multiple integrated text vectors will be obtained. For example, the integrated text vector of the vehicle service archive dimension is , the integrated text vector of the vehicle file dimension is , the integrated text vector of the driver profile dimension is , the integrated text vector of the violation file dimension is Next, these integrated text vectors are combined to obtain the bus usage representation vector. One way is to concatenate these vectors, that is, Another method is to perform a weighted combination based on the importance of each integrated text vector. The evaluation contribution coefficient of each integrated text vector can be determined based on information such as the vehicle weighting level, and then each integrated text vector is multiplied by the corresponding evaluation contribution coefficient and then added. Suppose the evaluation contribution coefficients are , then the bus usage characterization vector .

[0035] By performing vector integration processing on each first text sub-vector set and combining multiple integrated text vectors obtained by the integration, a bus usage representation vector that can comprehensively reflect the usage of the bus throughout its life cycle is obtained.

[0036] Step 500: Determine the usage specification evaluation result of the bus life cycle file through the bus usage characterization vector.

[0037] In step 500, the evaluation result of the use specification of the bus life cycle archive is determined by the bus use characterization vector. The bus use characterization vector is a vector representation that comprehensively reflects the use of the bus in the entire life cycle, and contains information in multiple dimensions such as vehicle service files, vehicle files, driver files, and violation files. In order to achieve this goal, a pre-trained model can be used, such as a second text processing network. After training with a large number of bus life cycle archive samples, the network can learn the mapping relationship between the bus use specifications and the bus use characterization vector. The obtained bus use characterization vector is input into the second text processing network, and the network will perform a full connection mapping on the vector. Full connection mapping means that each neuron in the network is connected to all neurons in the previous layer, and the input bus use characterization vector is converted into a probability distribution, i.e., the first evaluation probability distribution, through a series of linear transformations and nonlinear activation functions.

[0038] For example, assuming that the usage norms evaluation results are divided into three situations: "normative", "basic normative", and "non-normative", the first evaluation probability distribution output by the second text processing network after full-connection mapping may be [0.2, 0.3, 0.5], respectively representing the probabilities of the bus usage being "normative", "basic normative", and "non-normative".

[0039] The final evaluation result of the use specification is determined based on the first evaluation probability distribution. One method is to select the category with the largest probability as the evaluation result. In the above example, the probability of "non-standard" is the largest, which is 0.5, so the evaluation result of the use specification of the bus is determined to be "non-standard".

[0040] In specific implementation, the second text processing network can adopt a neural network structure such as a multi-layer perceptron (MLP). MLP consists of an input layer, a hidden layer, and an output layer. The input layer receives the bus usage representation vector, the hidden layer performs nonlinear transformation on the input through a series of neurons, and the output layer outputs the first evaluation probability distribution. Its calculation formula can be expressed as , where x is the bus usage characterization vector, and are the weight matrix and bias vector of the i-th layer, f is the activation function, such as ReLU, Sigmoid, etc., and y is the first evaluation probability distribution.

[0041] In this way, the bus usage representation vector can be used with the help of the second text processing network to accurately determine the usage specification evaluation results of the bus life cycle archive, providing an important basis for the management and supervision of the bus.

[0042] In one implementation, step 400, performing vector integration processing on each first text sub-vector set, and combining multiple integrated text vectors obtained by integration to obtain a bus usage representation vector of the bus life cycle archive, includes: Step 410: Determine a first text processing network corresponding to each first text sub-vector set; Step 420: Load the first text sub-vector set corresponding to each text embedding vector into the corresponding first text processing network for vector integration processing to obtain an integrated text vector corresponding to each text embedding vector; Step 430: Perform vector combination on the multiple integrated text vectors to obtain a bus usage representation vector of the bus life cycle file.

[0043] In step 410, the first text processing network corresponding to each first text sub-vector set is determined. The first text sub-vector sets of different archive record dimensions have different characteristics and information structures, and require targeted processing networks for effective feature extraction and integration. Taking the previously mentioned unit official vehicles as an example, the first text sub-vector set of the vehicle service archive dimension contains information on vehicle annual inspection, insurance, maintenance, etc., which are related in time series and business logic; the first text sub-vector set of the vehicle archive dimension is mainly the basic information of the vehicle, which is relatively static and structured; the first text sub-vector set of the driver archive dimension involves information such as the driver's qualifications and driving habits; the first text sub-vector set of the violation archive dimension records the vehicle's violations, including the time, location, type, etc. of the violation. According to the characteristics of these different dimensions, a suitable first text processing network will be assigned to each first text sub-vector set. These networks can be neural networks based on deep learning, such as the long short-term memory network (LSTM), which is suitable for processing the vehicle file dimension with time series characteristics; the convolutional neural network (CNN) may be more suitable for feature extraction of structured information in the vehicle file dimension; the Transformer architecture can be used in the driver file and violation file dimensions to capture the complex dependencies between information.

[0044] In step 420, the first text sub-vector set corresponding to each text embedding vector is loaded into the corresponding first text processing network for vector integration processing to obtain an integrated text vector corresponding to each text embedding vector. Taking the vehicle service file dimension as an example, assuming that its first text sub-vector set contains multiple sub-vectors reflecting the vehicle status at different time points, such as When using the LSTM network for processing, the LSTM network can process these sub-vectors sequentially through its internal memory units and gating mechanisms to capture long-term dependencies in the time series. The core formula of LSTM includes the input gate , Forget Gate , cell state update and output gate , where x t is the current input subvector, h t-1 is the hidden state of the previous moment, W is the weight matrix, b is the bias vector, is the Sigmoid function, It is element-by-element multiplication. After a series of processing, an integrated vector is finally output, that is, the integrated text vector of the vehicle file dimension. For the vehicle file dimension, if CNN is used for processing, CNN will slide the convolution kernel on the first text sub-vector set to extract local features, reduce the feature dimension through pooling operations, and finally obtain the integrated text vector.

[0045] In step 430, multiple integrated text vectors are combined to obtain a bus usage representation vector for the bus life cycle archive. The integrated text vector of each archive record dimension contains important information about the bus usage in that dimension, but the information of a single dimension is not enough to fully reflect the overall usage of the bus. Therefore, it is necessary to combine the integrated text vectors of the vehicle service archive dimension, vehicle archive dimension, driver archive dimension, and violation archive dimension. One method is the splicing operation, that is, connecting these integrated text vectors in a certain order. Assume that the integrated text vector of the vehicle service archive dimension is , the integrated text vector of the vehicle file dimension is , the integrated text vector of the driver profile dimension is , the integrated text vector of the violation file dimension is , then the bus usage characterization vector Another way is weighted summation. According to the importance of different file record dimensions in the evaluation of bus use, a corresponding weight is assigned to each integrated text vector, and then the weighted integrated text vectors are added to obtain the bus use representation vector. Suppose the weights of the vehicle file dimension, vehicle file dimension, driver file dimension, and violation file dimension are ,and , then the bus usage characterization vector .

[0046] Through steps 410-S430, the conversion from the first text sub-vector set of different archive record dimensions to the bus usage representation vector is completed. This process fully considers the characteristics and importance of information of different dimensions, integrates the scattered information into a comprehensive vector representation, and lays a solid foundation for the subsequent accurate evaluation of bus usage specifications. Different first text processing networks and vector combination methods can be selected and optimized according to actual conditions to adapt to different bus management needs and data characteristics.

[0047] In one implementation, step 430, combining multiple integrated text vectors to obtain a bus usage representation vector of the bus life cycle archive includes: Step 431: Determine the evaluation contribution coefficient of each integrated text vector; wherein the process of determining the evaluation contribution coefficient includes: obtaining the vehicle weighting level; determining the evaluation contribution coefficient of each integrated text vector based on the correlation information between each archive record dimension and the vehicle weighting level.

[0048] Step 432: performing weighted adjustment on each integrated text vector by evaluating the contribution coefficient to obtain a plurality of adjusted text vectors; Step 433: Perform vector combination on multiple adjustment text vectors to obtain a bus usage representation vector of the bus life cycle file.

[0049] First, obtain the vehicle authorization level. The vehicle authorization level includes information such as what positions can use it and the usage threshold. It reflects the importance and normative requirements of the use of official vehicles. Taking the official vehicles of the unit as an example, the official vehicles used by personnel at different levels may have different authorization levels. The official vehicles used by senior personnel often involve important government activities, so the authorization level is higher, and the requirements for vehicle maintenance, insurance, driver qualifications, etc. are more stringent; while the official vehicles used by ordinary staff have relatively low authorization levels, and the usage specifications and requirements may be relatively loose. The vehicle authorization level information can be obtained from the database of the official vehicle management system. This information can be stored in the form of structured data, such as a vehicle use permission table, which records the user position, usage scope, and other related information corresponding to each official vehicle.

[0050] In step 4302, the evaluation contribution coefficient of each integrated text vector is determined based on the correlation information between each file record dimension and the vehicle empowerment level. The evaluation contribution coefficient reflects the importance of each file record dimension in the evaluation of the bus use specifications. For buses with higher empowerment levels, the vehicle service file dimension (such as annual inspection, insurance, maintenance, etc.) and the driver file dimension (such as driver basic information related materials) are closely related to the normal use and safety of the vehicle, so the evaluation contribution coefficient of the integrated text vector of these two dimensions may be relatively high. For buses with lower empowerment levels, the violation file dimension (such as violation records, violators, punishment content, etc.) may account for a large proportion in the evaluation, because this type of bus focuses more on the normativeness of daily use. The evaluation contribution coefficient can be determined by establishing a correlation model, for example, using a linear regression model, with the vehicle empowerment level as the independent variable and the importance score of each file record dimension as the dependent variable. Through training with a large amount of historical data, a linear relationship between each file record dimension and the vehicle empowerment level is obtained, thereby calculating the evaluation contribution coefficient. Assume that the evaluation contribution coefficients of the vehicle service file dimension, vehicle file dimension, driver file dimension and violation file dimension are respectively ,and For a bus with a higher level of empowerment, it may be ; For a bus with a lower empowerment level, it may .

[0051] After determining the evaluation contribution coefficient, in step 431, it is clear that each integrated text vector has its corresponding evaluation contribution coefficient, which is determined based on the correlation between the vehicle weighting level and the file record dimension in the previous step. For example, the integrated text vector V of the vehicle service file dimension cs Corresponding evaluation contribution coefficient , the integrated text vector V of the vehicle profile dimension vehicle Corresponding evaluation contribution coefficient , the integrated text vector V of the driver profile dimension driver Corresponding evaluation contribution coefficient , the integrated text vector V of the violation file dimension violation Corresponding evaluation contribution coefficient .

[0052] In step 432, each integrated text vector is weighted and adjusted by evaluating the contribution coefficient to obtain multiple adjusted text vectors. The purpose of weighted adjustment is to highlight the importance of different file record dimensions in the evaluation of bus use. Each integrated text vector is multiplied by its corresponding evaluation contribution coefficient to obtain an adjusted text vector. Specifically, the adjusted text vector of the vehicle service file dimension is , the adjusted text vector of the vehicle file dimension is , the adjusted text vector of the driver profile dimension is , the adjusted text vector of the violation file dimension is Taking the buses with a higher weighting level as an example, the integrated text vector V of the vehicle service file dimension cs Multiplying by the higher assessment contribution factor After that, its influence in subsequent combinations will be relatively greater, because vehicle service files are crucial to the normal operation and safety of high-level buses.

[0053] In step 433, multiple adjustment text vectors are combined to obtain the bus usage representation vector of the bus life cycle archive. Vector combination is to integrate the information of each archive record dimension after weighted adjustment to form a comprehensive vector representation to fully reflect the usage of the bus. The combination method used is to add multiple adjustment text vectors, and the formula is In this way, the information of different file record dimensions is effectively integrated according to their importance. For example, for a bus with a higher empowerment level, the adjustment text vectors of the vehicle management file dimension and the driver file dimension have a greater impact on the final bus use representation vector during the addition process, making the vector more able to reflect the requirements of high-level buses in terms of vehicle management and driver qualifications; while for buses with a lower empowerment level, the adjustment text vector of the violation file dimension will occupy a more important position in the bus use representation vector due to its higher evaluation contribution coefficient, highlighting the importance of the standardization of the use of such buses.

[0054] In practice, matrix operations can be used to efficiently complete these steps. For the evaluation contribution coefficient, it can be stored as a coefficient vector , the integrated text vector can be stored as a matrix , then the bus usage characterization vector This matrix operation method can make full use of the parallel computing capabilities of computers and improve processing efficiency.

[0055] Through steps 431-S433, the influence of vehicle weighting levels on different file record dimensions is fully considered, and the integrated text vectors are reasonably weighted and combined to obtain a bus usage representation vector that can accurately reflect the bus usage.

[0056] In one implementation, after extracting archival text segments of multiple archival record dimensions from the historical archival record text in step 200 and mapping the extracted multiple archival text segments into multiple text embedding vectors, the method further includes: Step 210: Using the third vector coverage as the scale size and the fourth vector coverage as the scale step, perform mobile scaling on each text embedding vector to obtain a second text sub-vector set corresponding to each text embedding vector. Based on this, as another implementation scheme of step 400, that is, perform vector integration processing on each first text sub-vector set, and combine the multiple integrated text vectors obtained by integration to obtain the bus usage representation vector of the bus life cycle archive, including: Step 401: performing vector integration processing on each first text sub-vector set, and combining multiple first integrated sub-text vectors obtained by integration to obtain a first bus use sub-representation vector, and performing vector integration processing on each second text sub-vector set, and combining multiple second integrated sub-text vectors obtained by integration to obtain a second bus use sub-representation vector; Step 402: Combine the first bus usage sub-characterization vector and the second bus usage sub-characterization vector to obtain a bus usage characterization vector of the bus life cycle file.

[0057] In step 210, after the historical archive record text is extracted into archive text fragments of multiple archive record dimensions and these fragments are mapped into multiple text embedding vectors, each text embedding vector is subjected to mobile scaling again. This time, the third vector coverage is used as the scaling size, and the fourth vector coverage is used as the scaling stride, so as to obtain a second text sub-vector set corresponding to each text embedding vector. It can be understood that the coverage ranges of each vector in the embodiment of the present application can be pre-set according to actual needs, and there is no specific limitation. The purpose of step 210 is to extract features from text embedding vectors at different scales and granularities to capture more information at different levels. Taking the official vehicles of the unit mentioned above as an example, it is assumed that a text embedding vector of the vehicle service archive dimension contains a vector representation of information such as vehicle annual inspection, insurance, and maintenance. The first mobile scaling process (corresponding to step 300) may extract some features at a specific scale, and in step 210, different scaling sizes and strides are used for reprocessing. For example, the first time the scale is larger, some macro information can be extracted, while this time a smaller scale and stride are used, more detailed local features can be mined, such as a more accurate representation of the specific time and cost of a vehicle maintenance in the vector. The method for implementing this mobile scale processing is similar to step 300, and the text embedding vector can be looped through and the interception operation is performed according to the set scale and stride, thereby obtaining a second text sub-vector set.

[0058] In step 401, vector integration processing is performed on each first text sub-vector set, and multiple first integrated sub-text vectors obtained by integration are combined to obtain a first bus use sub-representation vector. At the same time, vector integration processing is performed on each second text sub-vector set, and multiple second integrated sub-text vectors obtained by integration are combined to obtain a second bus use sub-representation vector. For the vector integration processing of the first text sub-vector set, methods such as long short-term memory network (LSTM), gated recurrent unit (GRU) or attention mechanism mentioned above can be used. Taking LSTM as an example, for the first text sub-vector set of the vehicle service file dimension, the LSTM network will process the input sub-vector sequence according to its internal memory unit and gating mechanism, and output an integrated vector, that is, the first integrated sub-text vector. The first integrated sub-text vectors of multiple archive record dimensions are combined by splicing or weighted summation to obtain the first bus use sub-representation vector. Similarly, for the second text sub-vector set, the same or different vector integration methods are also used for processing to obtain the second integrated sub-text vector and combine it into the second bus use sub-representation vector. Assume that the first integrated subtext vectors of the vehicle service file dimension, vehicle file dimension, driver file dimension, and violation file dimension are , the first bus usage sub-characterization vector is obtained by weighted summation ,in is the corresponding weight coefficient. Similarly, for the second integrated sub-text vector , combined to get the second bus usage sub-characterization vector .

[0059] In step 402, the first bus usage sub-characterization vector and the second bus usage sub-characterization vector are combined to obtain a bus usage characterization vector of the bus life cycle file. This combination process can further integrate feature information of different scales and levels, making the final bus usage characterization vector more comprehensive and accurate. The combination method can be splicing or weighted summation.

[0060] Through step 210 and steps 401-402, the text information of the bus life cycle archive is feature extracted and integrated from different scales and levels, and finally a more comprehensive and accurate bus usage representation vector is generated. This vector can better reflect the usage of the bus throughout its life cycle, provide a basis for subsequent bus usage specification evaluation, help improve the accuracy and reliability of the evaluation, and meet the actual needs of bus management.

[0061] In one implementation, step 500, determining the use specification evaluation result of the bus life cycle archive through the bus use characterization vector, includes: Step 510: Load the bus usage representation vector into the second text processing network for full connection mapping to obtain a first evaluation probability distribution; Step 520: Determine the evaluation result of the usage specification of the bus life cycle file according to the first evaluation probability distribution.

[0062] In step 510, the bus usage representation vector is loaded into the second text processing network for full connection mapping to obtain a first evaluation probability distribution. The second text processing network is a pre-trained neural network model, such as a multi-layer perceptron (MLP). The multi-layer perceptron consists of an input layer, a hidden layer, and an output layer. The input layer receives the bus usage representation vector, the hidden layer contains multiple neurons, and processes the input through a series of linear transformations and nonlinear activation functions. The output layer outputs the first evaluation probability distribution.

[0063] Assume that the bus usage characterization vector is V bur , whose dimension is n. The input layer of the second text processing network has n neurons, corresponding to the dimension of the bus usage representation vector. The number of neurons in the hidden layer can be adjusted according to the actual situation. Suppose the hidden layer has m neurons. The linear transformation from the input layer to the hidden layer can be expressed as , where W1 is an m×n weight matrix and b_1 is an m-dimensional bias vector. Then, Z1 is processed by a nonlinear activation function f (such as ReLU function, f(x)=max(0,x)) to obtain the output H=f(Z1) of the hidden layer.

[0064] The linear transformation from the hidden layer to the output layer is , where W2 is a k×m weight matrix (k is the number of neurons in the output layer, corresponding to the number of categories of the evaluation results), and b2 is a k-dimensional bias vector. Finally, Z2 is converted into a probability distribution through the Softmax function, that is, the first evaluation probability distribution P=Softmax(Z2). The formula of the Softmax function is , where z i is the i-th element of Z2.

[0065] Assuming that the evaluation results of bus usage regulations are divided into three situations: "standard", "basic standard" and "non-standard", then the output layer has 3 neurons, and the first evaluation probability distribution P=[p1, p2, p3], which respectively represent the probabilities of bus usage being "standard", "basic standard" and "non-standard", and p1+p2+p3=1.

[0066] In step 520, the evaluation result of the use specification of the bus life cycle file is determined according to the first evaluation probability distribution. One method is to select the category with the highest probability as the evaluation result. For example, if the first evaluation probability distribution is [0.2, 0.3, 0.5], then the probability of "non-standard" is the highest, and the evaluation result of the use specification of the bus is determined to be "non-standard".

[0067] To achieve this process, the Argmax function can be used, that is, Result = argmax (P), where Result is the index of the evaluation result. If Result = 0, the evaluation result is "standard"; if Result = 1, the evaluation result is "basic standard"; if Result = 2, the evaluation result is "non-standard".

[0068] In practical applications, the training of the second text processing network is based on a large number of bus life cycle archive samples. By adjusting the weight matrices W1, W2 and the bias vectors b1, b2, the network can accurately map the bus usage representation vector to the corresponding usage specification evaluation results. The training process uses a loss function (such as the cross entropy loss function) to measure the difference between the predicted results and the true labels, and updates the network parameters through the back-propagation algorithm.

[0069] In one implementation, multiple first text processing networks and second text processing networks are obtained by parallel training of bus life cycle archive samples, and multiple first text processing networks and second text processing networks are trained through the following process: Step 10: Obtaining archive record training text, which includes full-cycle archive record training text of multiple bus training samples and supervision marks of usage specification evaluation results of each bus training sample; Step 20: extracting archive training text segments of multiple archive record dimensions from the full-cycle archive record training text, and mapping the extracted multiple archive training text segments into multiple training text embedding vectors; Step 30: Using the first vector coverage as the scaling size and the second vector coverage as the scaling step, perform mobile scaling on each training text embedding vector to obtain a first training text sub-vector set corresponding to each training text embedding vector; Step 40: Load the first training text sub-vector set corresponding to each training text embedding vector into the corresponding first text processing network for vector integration processing, and combine the training integrated text vectors corresponding to each training text embedding vector obtained by integration to obtain a bus sample usage representation vector; Step 50: Load the bus sample usage representation vector into the second text processing network for full connection mapping to obtain a second evaluation probability distribution, and determine the mapping error value through the second evaluation probability distribution and the corresponding usage specification evaluation result supervision mark; Step 60: Convergence training is performed on the network parameters of the plurality of first text processing networks and the second text processing networks by mapping the error values.

[0070] In step 10, the archive record training text is obtained, which includes the full-cycle archive record training text of multiple bus training examples and the supervision mark of the use specification evaluation result of each bus training example. The bus training example is a representative sample selected from the actual use data of a large number of buses. The full-cycle archive record training text covers various information from the purchase, use to scrapping of the bus, such as annual inspection, insurance, and maintenance records in the vehicle archive dimension, basic vehicle information in the vehicle archive dimension, driver qualifications and driving records in the driver archive dimension, and violations in the violation archive dimension. The supervision mark of the use specification evaluation result is a clear mark of whether the use of each bus training example is standardized, for example, it can be represented by labels such as "standard", "basic standard", and "non-standard". Taking the official vehicles of the unit as an example, the full-cycle archive records of 1,000 buses may be collected from the database as training texts, and the use specification evaluation results are marked for each bus, such as the first vehicle is marked as "standard", the second vehicle is marked as "non-standard", etc.

[0071] In step 20, multiple archive training text segments of archive record dimensions are extracted from the full-cycle archive record training text, and the extracted multiple archive training text segments are mapped into multiple training text embedding vectors. The archive training text segments can be extracted using a rule-based method or a machine learning method. The rule-based method locates text segments of different archive record dimensions by defining keyword matching rules or regular expressions, for example, matching sentences containing keywords such as "annual inspection" and "insurance" as text segments of vehicle service archive dimensions. Machine learning methods such as named entity recognition (NER) models can learn entity information of different archive record dimensions through training, so as to extract text segments more accurately. When mapping the archive training text segments into training text embedding vectors, word vector models (such as Word2Vec, GloVe) or deep learning-based language models (such as BERT, XLNet) can be used. Taking Word2Vec as an example, it maps each word to a vector space of fixed dimension. For an archive training text segment, each word therein is converted into a corresponding word vector, and then these word vectors are combined into a training text embedding vector by averaging or summing. Assume that an archive training text segment is "Vehicle annual inspection passed", convert the three words "vehicle", "annual inspection" and "passed" into word vectors v1, v2 and v3 respectively, and then calculate the training text embedding vector V=(v1+v2+v3) / 3.

[0072] In step 30, the first vector coverage range is used as the scale size, and the second vector coverage range is used as the scale step, and each training text embedding vector is subjected to mobile scaling processing to obtain a first training text sub-vector set corresponding to each training text embedding vector. Mobile scaling processing is a method for extracting local features on sequence data. By setting the scale size and the scale step, a window is slid on the training text embedding vector, and each time the vector in the window is intercepted as a sub-vector. For examples, please refer to the aforementioned step 300, which will not be described in detail here.

[0073] In step 40, the first training text sub-vector set corresponding to each training text embedding vector is loaded into the corresponding first text processing network for vector integration processing, and the training integrated text vectors corresponding to each training text embedding vector obtained by integration are combined to obtain the bus sample usage representation vector. The first text processing network can be a long short-term memory network (LSTM), a gated recurrent unit (GRU) or a Transformer architecture. Taking LSTM as an example, it can process long-term dependencies in sequence data through its internal memory units and gating mechanisms. After being processed by the LSTM network, an integrated vector is finally output, namely the training integrated text vector. The training integrated text vectors of multiple archive record dimensions are combined, such as by splicing or weighted summation to obtain the bus sample usage representation vector.

[0074] In step 50, the bus sample usage representation vector is loaded into the second text processing network for full connection mapping to obtain a second evaluation probability distribution, and the mapping error value is determined by the second evaluation probability distribution and the corresponding usage specification evaluation result supervision mark. The second text processing network can be a multi-layer perceptron (MLP), which consists of an input layer, a hidden layer, and an output layer. The input layer receives the bus sample usage representation vector, the hidden layer processes the input through a series of linear transformations and nonlinear activation functions, and the output layer outputs the second evaluation probability distribution. Assume that the bus usage specification evaluation results are divided into three cases: "standard", "basic standard", and "non-standard", and the output layer has 3 neurons. The second evaluation probability distribution P=[p1, p2, p3] represents the probability that the bus usage is "standard", "basic standard", and "non-standard". The cross entropy loss function is used to measure the difference between the second evaluation probability distribution and the usage specification evaluation result supervision mark. The formula of the cross entropy loss function is ,in is the ith element labeled with the canonical evaluation result supervision (if it is “canonical” then ; If it is a "basic specification" then If it is "non-standard", ), p iis the i-th element of the second evaluation probability distribution, and k is the number of categories of the evaluation results. This loss value is the mapping error value, which reflects the degree of difference between the network prediction result and the true label.

[0075] In step 60, the network parameters of the first text processing network and the second text processing network are trained to converge by mapping the error value. The back propagation algorithm is used to calculate the gradient of the mapping error value to the network parameters (such as the weight matrix and the bias vector), and then the parameters are updated according to the gradient. The feasible optimization algorithms include stochastic gradient descent (SGD), Adagrad, Adadelta, Adam, etc. Taking the stochastic gradient descent algorithm as an example, its update formula is ,in represents network parameters (such as weight matrix W or bias vector b), is the learning rate, which controls the step size of parameter update. It is the gradient of the mapping error value to the parameter variable. Through multiple iterations of training, the network parameters are continuously adjusted to gradually reduce the mapping error value until the convergence condition is reached (such as the error value is less than a certain threshold or the number of iterations reaches the upper limit). At this time, the network performance is optimal and can accurately extract features from the bus life cycle archive and evaluate the bus usage specifications.

[0076] In one implementation, after extracting archive training text segments of multiple archive record dimensions from the full-cycle archive record training text in step 20, and mapping the extracted multiple archive training text segments into multiple training text embedding vectors, the method further includes: Step 201: Using the third vector coverage as the scale size and the fourth vector coverage as the scale stride, each training text embedding vector is subjected to mobile scaling processing to obtain a second training text sub-vector set corresponding to each training text embedding vector; based on this, step 40: loading the first training text sub-vector set corresponding to each training text embedding vector into the corresponding first text processing network for vector integration processing, and combining the training integrated text vectors corresponding to each training text embedding vector obtained by integration to obtain a bus sample usage representation vector, including: Step 41: Load the first training text sub-vector set corresponding to each training text embedding vector into the corresponding first text processing network for vector integration processing, and combine the first training integrated sub-text vectors corresponding to each training text embedding vector obtained by integration to obtain a first training bus usage sub-representation vector; Step 42: Load the second training text sub-vector set corresponding to each training text embedding vector into the corresponding first text processing network for vector integration processing, and combine the second training integrated sub-text vectors corresponding to each training text embedding vector obtained by integration to obtain a second training bus usage sub-representation vector; Step 43: Combine the first training bus usage sub-characterization vector and the second training bus usage sub-characterization vector to obtain a bus sample usage characterization vector of the bus training sample.

[0077] Step 20 completes the extraction of archive training text segments of multiple archive record dimensions in the full-cycle archive record training text, and maps them into multiple training text embedding vectors. On this basis, in step 201, the third vector coverage is used as the scale size, and the fourth vector coverage is used as the scale step, and each training text embedding vector is subjected to mobile scaling processing to obtain a second training text sub-vector set corresponding to each training text embedding vector. The purpose of this step is to extract features from the training text embedding vector from different scales and granularities. Taking the archive records of the unit's official vehicles mentioned above as an example, it is assumed that a training text embedding vector of the vehicle archive dimension contains the vehicle's annual inspection, insurance and maintenance information for many years. In step 30, a mobile scaling process has been performed according to the first vector coverage and the second vector coverage, and the first training text sub-vector set has been obtained. In step 201, different scaling sizes and strides are used for reprocessing. For example, the first processing with a larger scale may extract features about the overall vehicle service status of a vehicle in a certain year; this time, a smaller scale and stride can be used to mine more detailed information, such as the detailed cost of a specific repair and a more accurate representation of the repair items in the vector. By looping through the training text embedding vector, the interception operation is performed according to the set third vector coverage range and fourth vector coverage range, thereby obtaining the second training text sub-vector set.

[0078] Based on the second training text sub-vector set obtained in step 201, in step 41, the first training text sub-vector set corresponding to each training text embedding vector is loaded into the corresponding first text processing network for vector integration processing, and the first training integrated sub-text vectors corresponding to each training text embedding vector obtained by integration are combined to obtain the first training bus use sub-representation vector. The first text processing network can be a long short-term memory network (LSTM), a gated recurrent unit (GRU) or a Transformer architecture. After a series of processing, an integrated vector is output, namely the first training integrated sub-text vector. For multiple file record dimensions, such as the vehicle service file dimension, the vehicle file dimension, the driver file dimension and the violation file dimension, the first training integrated sub-text vector of each dimension is combined. The combination method can be splicing or weighted summation.

[0079] In step 42, the second training text sub-vector set corresponding to each training text embedding vector is loaded into the corresponding first text processing network for vector integration processing, and the second training integrated sub-text vectors corresponding to each training text embedding vector obtained by integration are combined to obtain the second training bus usage sub-representation vector. Similar to step 41, the second training text sub-vector set is processed using the same or different first text processing network. Since the second training text sub-vector set is obtained from a different scale, it contains feature information at a different level from the first training text sub-vector set. For example, for the vehicle file dimension, the second training text sub-vector set may extract more detailed features of the vehicle configuration details.

[0080] In step 43, the first training bus usage sub-characterization vector and the second training bus usage sub-characterization vector are combined to obtain a bus sample usage characterization vector of the bus training sample. This step integrates feature information of different scales and levels so that the final bus sample usage characterization vector can more comprehensively and accurately reflect the usage of the bus. The combination method can also be concatenation or weighted summation.

[0081] Through step 201 and steps 41 - S43, the training text embedding vectors of the bus life cycle archive are feature extracted and integrated from different scales and levels, and finally a more comprehensive and accurate bus sample usage representation vector is generated. This vector can better reflect the use of the bus throughout its life cycle, provide a stronger basis for the subsequent bus use specification evaluation, help improve the accuracy and reliability of the evaluation, and meet the actual needs of bus management.

[0082] In one implementation, each first text processing network includes a text precoding component and a vector integration component. At this time, as another implementation of step 40, that is, loading the first training text sub-vector set corresponding to each training text embedding vector into the corresponding first text processing network for vector integration processing, and combining the training integrated text vectors corresponding to each training text embedding vector obtained by integration to obtain the bus sample usage representation vector, including: Step 4010: Perform text precoding on a first training text sub-vector set corresponding to each training text embedding vector through a text precoding component to obtain a first training precoded vector; Step 4020: performing vector integration processing on the first training precoding vector corresponding to each training text embedding vector through a vector integration component to obtain a training integrated text vector corresponding to each training text embedding vector; Step 4030: Combine the training integrated text vectors corresponding to each training text embedding vector to obtain a bus sample usage representation vector.

[0083] In step 4010, the first training text sub-vector set corresponding to each training text embedding vector is text pre-coded by the text pre-coding component to obtain a first training pre-coded vector. The purpose of text pre-coding is to map the first training text sub-vector set to a latent space dimension that is more suitable for subsequent processing. Taking the previously mentioned archive records of the unit's official vehicles as an example, it is assumed that the first training text sub-vector set of the vehicle archive dimension contains vector representations of relevant information such as vehicle annual inspection, insurance, and maintenance. The text pre-coding component can use linear projection to project each sub-vector in the first training text sub-vector set to the latent space dimension of the Transformer. The formula for linear projection can be expressed as V pe =W× V sv +b, where V sv is the sub-vector in the first training text sub-vector set, W is the weight matrix, b is the bias vector, V pe is the first training pre-coded vector. Through this linear transformation, the original sub-vector can be converted into a pre-coded vector with richer feature representation, which is convenient for subsequent vector integration components to process.

[0084] In step 4020, the first training pre-coding vector corresponding to each training text embedding vector is vector-integrated by the vector integration component to obtain a training integrated text vector corresponding to each training text embedding vector. The function of the vector integration component is to further process and integrate the pre-coding vector to extract more representative features. The vector integration component can adopt a variety of architectures, such as the multi-head attention mechanism in the Transformer architecture. The multi-head attention mechanism enables the model to focus on different parts of the input in different representation subspaces, thereby more comprehensively capturing the relationship between the input vectors. The formula of the multi-head attention mechanism is ,in , Q, K, V are query, key and value matrices respectively, and is a learnable weight matrix. The first training pre-coded vector is taken as input, the attention weights of different parts are calculated through the multi-head attention mechanism, and then the vector is weighted summed to obtain the integrated vector, that is, the training integrated text vector.

[0085] In step 4030, the training integrated text vectors corresponding to each training text embedding vector are combined to obtain a bus sample usage representation vector. The training integrated text vectors of different archive record dimensions contain important information about bus usage in their respective dimensions, and this information needs to be integrated. The combination method can be concatenation or weighted summation.

[0086] Through steps 4010-S4030, the text precoding component and vector integration component in the first text processing network are used to effectively process and integrate the first training text sub-vector set corresponding to each training text embedding vector. The text precoding component projects the original sub-vector to a latent space dimension that is more suitable for processing, and the vector integration component further extracts and integrates features, and finally combines the training integrated text vectors of different dimensions into a bus sample usage representation vector. This process makes full use of the structure and function of the first text processing network, providing a more accurate and representative input vector for the subsequent evaluation of bus usage specifications.

[0087] In one implementation, each first text processing network further includes a position embedding component. In step 4010, a text precoding component is used to perform text precoding on a first training text sub-vector set corresponding to each training text embedding vector. After obtaining the first training precoded vector, the method further includes: Step 4011: position embedding is performed on the first training text sub-vector set corresponding to each training text embedding vector by the position embedding component to obtain the first training text position sub-vector set. Based on this, step 4020, vector integration processing is performed on the first training pre-coding vector corresponding to each training text embedding vector by the vector integration component to obtain the training integration text vector corresponding to each training text embedding vector, including: Step 4021: Perform vector integration processing on the first training text position sub-vector set corresponding to each training text embedding vector through the vector integration component to obtain a training integrated text vector corresponding to each training text embedding vector.

[0088] Step 4011 after step 4010 and implementation scheme of step 4020 Step 4021 is an important step for further refining the operation of the first text processing network when processing the bus life cycle archive training data, with the purpose of better integrating information and generating a more representative bus sample usage representation vector.

[0089] In step 4010, the first training text sub-vector set corresponding to each training text embedding vector has been text pre-encoded by the text pre-encoding component to obtain the first training pre-encoded vector. On this basis, in step 4011, the first training text sub-vector set corresponding to each training text embedding vector is positionally embedded by the position embedding component to obtain the first training text position sub-vector set. The role of position embedding is to add position information to the input sub-vector set, because when processing sequence data, the order and position of the vectors often contain important semantic information. For example, in the vehicle service file dimension of a public bus, the order of information such as the vehicle's annual inspection time and insurance purchase time is very critical for judging the vehicle's usage specifications and maintenance conditions.

[0090] In order to achieve position embedding, we first determine the importance of the archive dimension corresponding to each set of training text sub-vectors. The importance of the archive dimension can be set according to the role of different archive record dimensions in the evaluation of the use of public vehicles. For example, for a unit's official vehicle, the vehicle archive dimension (including annual inspection, insurance, maintenance, etc.) and the violation archive dimension (including violation records, penalty content, etc.) are relatively more important in evaluating whether the use of public vehicles is standardized, so a higher priority can be set for these two dimensions. The importance of the vehicle archive dimension (mainly basic vehicle information) and the driver archive dimension (basic driver information) may be relatively low.

[0091] After determining the importance of the archive dimension, the first training text sub-vector set corresponding to each training text embedding vector is positionally embedded using the importance. One position embedding method is to use a combination of sine and cosine functions. Assuming that the sub-vector dimension in the first training text sub-vector set is d, the position index is pos, and the dimension index is i, then the position embedding vector . Add these position embedding vectors to the sub-vectors in the first training text sub-vector set to obtain the first training text position sub-vector set. In this way, each sub-vector contains its position information in the sequence, which helps the subsequent model better understand the order and structure of the data.

[0092] In step 4021, the vector integration component performs vector integration processing on the first training text position sub-vector set corresponding to each training text embedding vector to obtain the training integrated text vector corresponding to each training text embedding vector. The function of the vector integration component is to integrate the sub-vector set after position embedding to extract more representative features. The vector integration component can include multiple sub-components, such as an internal attention component, a first cross-layer identity mapping component and a normalization component, a multi-layer perceptron component, and a second cross-layer identity mapping component and a normalization component.

[0093] The internal attention component performs internal weighted focusing on the first training text position sub-vector set corresponding to each training text embedding vector to obtain the first focused encoding vector. The internal attention mechanism allows the model to focus on the importance of different parts in the sub-vector set. By calculating the correlation between each sub-vector and other sub-vectors, a weight is assigned to each sub-vector, and then all sub-vectors are weighted summed according to these weights to obtain the first focused encoding vector. For example, in the vehicle service file dimension, if several sub-vectors contain key annual inspection information, the internal attention mechanism will assign higher weights to these sub-vectors, thereby highlighting these important information during the integration process.

[0094] Next, the first cross-layer identity mapping component and the standardization component perform standardization operations on the first focused encoding vector and the first training pre-encoding vector corresponding to each training text embedding vector to obtain the first standardized encoding vector. The role of the cross-layer identity mapping (ResNet) is to solve the gradient vanishing and gradient exploding problems in deep networks. By adding the input directly to the output, the network can learn more complex features. Standardization operations (such as Layer Normalization) can normalize the input vector to make the model training more stable. Specifically, the formula for LayerNormalization is ,in and are the mean and variance of the input vector, is a small constant, and are learnable parameters.

[0095] Then, the multi-layer perceptron component performs forward propagation on the first standardized encoding vector corresponding to each training text embedding vector to obtain the first forward propagation result. The multi-layer perceptron is a neural network with multiple hidden layers. It processes the input through a series of linear transformations and nonlinear activation functions (such as the ReLU function) to extract more advanced features.

[0096] Finally, the second cross-layer identity mapping component and the standardization component perform standardization operations on the first standardized encoding vector and the first forward propagation result corresponding to each training text embedding vector to obtain the training integrated text vector corresponding to each training text embedding vector. This step again uses cross-layer identity mapping and standardization operations to further optimize the feature representation, so that the final training integrated text vector can more accurately reflect the use of the bus in the dimension of the archive record.

[0097] Through step 4011 and step 4021, when processing the first training text sub-vector set, not only the semantic information of the text is considered, but also the position information is added, and the information is effectively integrated and optimized through the vector integration component. The training integrated text vector generated in this way can more comprehensively and accurately reflect the usage characteristics of the bus, providing more powerful support for the subsequent generation of bus sample usage representation vectors and the evaluation of bus usage specifications.

[0098] In one implementation, step 4011, position embedding is performed on a first training text sub-vector set corresponding to each training text embedding vector by a position embedding component to obtain a first training text position sub-vector set, including: Step 40111: Determine the importance of the archive dimension corresponding to each training text sub-vector set; Step 40112: Perform position embedding on the first training text sub-vector set corresponding to each training text embedding vector by using the archive dimension importance to obtain the first training text position sub-vector set.

[0099] In step 40111, the importance of the archive dimension corresponding to each training text sub-vector set is determined. The importance of the archive dimension reflects the relative importance of different archive record dimensions in the evaluation of the use of public vehicles. Taking the official vehicles of the unit as an example, different archive record dimensions, such as the vehicle archive dimension, the vehicle archive dimension, the driver archive dimension and the violation archive dimension, have different roles in evaluating whether the use of public vehicles is standardized. The vehicle archive dimension contains information such as the annual inspection, insurance, and maintenance of the vehicle. This information is directly related to the safety and legality of the vehicle and is crucial to the normal use of the public vehicle. For example, if the vehicle is not inspected on time or the insurance expires, it may face legal risks and safety hazards. Therefore, the vehicle archive dimension has a high importance in the evaluation. The vehicle archive dimension mainly records the basic information of the vehicle, such as the vehicle brand, model, frame number, etc. Although this information is the basic attribute of the public vehicle, it has a relatively small direct role in evaluating the use of public vehicles, so its importance is relatively low. The driver archive dimension involves the basic information, driving qualifications and driving records of the driver. The quality and behavior of the driver directly affect the safety and standardization of the use of the public vehicle, so this dimension also has a certain importance. The violation file dimension records the vehicle's violations, including the time, location, type and penalty of the violation. The violation record is an important basis for evaluating whether the use of public buses is standardized, so it is of high importance.

[0100] One way to determine the importance of profile dimensions is to set fixed priorities based on experience and expert knowledge. For example, set the priority of the vehicle service profile dimension to 80%, the priority of the violation profile dimension to 70%, the priority of the driver profile dimension to 60%, and the priority of the vehicle profile dimension to 30%. Another way is to determine the importance through data analysis. A large amount of bus usage data and the corresponding usage specification evaluation results can be analyzed to statistically analyze the correlation between the information of each profile dimension and the evaluation results. The higher the correlation, the greater the importance of the profile dimension in the evaluation. For example, the linear correlation between each profile dimension and the bus usage specification evaluation results is measured by calculating the Pearson correlation coefficient. The larger the absolute value of the correlation coefficient, the higher the importance of the profile dimension.

[0101] In step 40112, the first training text sub-vector set corresponding to each training text embedding vector is positionally embedded by using the archive dimension importance to obtain the first training text position sub-vector set. The purpose of position embedding is to add position information to the input sub-vector set, because when processing sequence data, the order and position of the vector often contain important semantic information. Position embedding can be achieved by using position encoding. One position encoding method is to use a combination of sine and cosine functions. Assuming that the sub-vector dimension in the first training text sub-vector set is d, the position index is pos, and the dimension index is i, the calculation formula of the position embedding vector is: For even dimensions (2i): ; For odd dimensions (2i+1): .

[0102] When performing position embedding, the position encoding will be adjusted according to the importance of the file dimension. For example, for file dimensions with higher importance, such as vehicle service file dimensions and violation file dimensions, the amplitude of the position encoding can be increased so that the sub-vectors of these dimensions are more easily noticed by the model in subsequent processing. Specifically, the position encoding of the file dimensions with higher importance can be multiplied by a coefficient greater than 1, such as 1.2 or 1.5. For file dimensions with lower importance, such as vehicle file dimensions, the position encoding can be multiplied by a coefficient less than 1, such as 0.8 or 0.5. In this way, the position information of important file dimensions can be highlighted and the model's ability to capture key information can be improved.

[0103] The adjusted position encoding vector is added to the sub-vectors in the first training text sub-vector set to obtain the first training text position sub-vector set. In this way, each sub-vector contains its position information in the sequence and the importance information of the archive dimension. In the subsequent vector integration processing, the model can better understand the order and importance of different archive dimension information based on this information, so as to more accurately extract features and generate a bus sample usage representation vector that can fully reflect the bus usage.

[0104] Through steps 40111-40112, the importance of different archival dimensions is fully considered in the position embedding process, and targeted position information is added to the first training text sub-vector set. This not only helps the model to better process sequence data, but also highlights the information of key archival dimensions, improves the effect of vector integration processing, and provides more accurate and representative input for the evaluation of bus use regulations.

[0105] In one implementation, the vector integration component includes an internal attention component, a first cross-layer identity mapping component and a normalization component, a multi-layer perceptron component, a second cross-layer identity mapping component and a normalization component. Step 4020, performing vector integration processing on the first training precoding vector corresponding to each training text embedding vector by the vector integration component to obtain a training integrated text vector corresponding to each training text embedding vector, including: Step 40201: Perform internal weight focusing on the first training pre-coding vector corresponding to each training text embedding vector through an internal attention component to obtain a first focused coding vector; Step 40202: performing a standardization operation on the first focused encoding vector and the first training pre-encoding vector corresponding to each training text embedding vector through a first cross-layer identity mapping component and a standardization component to obtain a first standardized encoding vector; Step 40203: forward propagating the first standardized encoding vector corresponding to each training text embedding vector through the multi-layer perceptron component to obtain a first forward propagation result; Step 40204: Perform a standardization operation on the first standardized encoding vector and the first forward propagation result corresponding to each training text embedding vector through the second cross-layer identity mapping component and the standardization component to obtain a training integrated text vector corresponding to each training text embedding vector.

[0106] Steps 40201 - 40204 in the implementation of step 4020 describe the detailed process of performing vector integration processing on the first training precoding vector corresponding to each training text embedding vector through the vector integration component to obtain the training integrated text vector corresponding to each training text embedding vector. The vector integration component includes an internal attention component, a first cross-layer identity mapping component and a normalization component, a multi-layer perceptron component, a second cross-layer identity mapping component and a normalization component, and these components work together to gradually process and optimize the input first training precoding vector, so that the final training integrated text vector can more accurately reflect the usage characteristics of the bus under various file record dimensions.

[0107] In step 40201, the first training pre-coding vector corresponding to each training text embedding vector is internally weighted focused by the internal attention component to obtain the first focused coding vector. The internal attention mechanism is a technology that enables the model to automatically focus on the importance of different parts of the input sequence. In the scenario of full life cycle safety management of buses, the first training pre-coding vector contains information from multiple archive record dimensions extracted from the bus life cycle archive, such as annual inspection, insurance, and maintenance information in the vehicle service archive dimension, basic vehicle information in the vehicle archive dimension, driver qualification information in the driver archive dimension, and violation record information in the violation archive dimension. This information may have different importance in different usage specification evaluation scenarios. For example, when evaluating whether a bus meets the safety driving standards, the annual inspection and maintenance information in the vehicle service archive dimension may be more critical; while when evaluating the compliance of bus use, the information in the violation archive dimension is more important.

[0108] The internal attention mechanism assigns a weight to each vector by calculating the correlation between each first training pre-coded vector and other vectors. Specifically, the internal attention mechanism first projects the first training pre-coded vector into the query, key, and value spaces, respectively, to obtain the query matrix Q, key matrix K, and value matrix V. Then, by calculating the dot product of the query matrix and the key matrix, the attention score matrix S=QK is obtained. T In order to prevent the dot product result from being too large, S can be scaled, that is , where d_k is the dimension of the key vector. Next, the attention score matrix is ​​converted into the attention weight matrix A=Softmax(S) using the Softmax function, so that the weight of each vector is between [0, 1] and the sum of the weights of all vectors is 1. Finally, the attention weight matrix is ​​multiplied by the value matrix to obtain the first focused encoding vector F=AV. In this way, the internal attention mechanism can automatically focus on the part of the input vector that is relevant to the current task, highlighting important information, thereby obtaining a more representative first focused encoding vector.

[0109] In step 40202, the first focused encoding vector and the first training pre-encoding vector corresponding to each training text embedding vector are standardized by the first cross-layer identity mapping component and the standardization component to obtain a first standardized encoding vector. Cross-layer identity mapping (ResNet) is a technology used to solve the gradient vanishing and gradient exploding problems in deep neural networks. During the vector integration process, as the number of network layers increases, the gradient may become very small or very large during the back propagation process, making the model difficult to train. The cross-layer identity mapping alleviates the gradient vanishing and gradient exploding problems by adding the input directly to the output so that the gradient can be propagated directly from the output layer to the input layer. Specifically, the first cross-layer identity mapping component adds the first training pre-encoding vector X to the first focused encoding vector F to obtain R=X+F.

[0110] The normalization component normalizes the result after addition to improve the training stability of the model. The normalization method is, for example, Layer Normalization, which has been introduced above and will not be repeated here. Through LayerNormalization, each dimension of the input vector is normalized to the same scale, so that the model can maintain a stable training effect under different inputs. After being processed by the first cross-layer identity mapping component and the normalization component, the first standardized encoding vector N1=LN(R) is obtained.

[0111] In step 40203, the first standardized encoding vector corresponding to each training text embedding vector is forward propagated through the multi-layer perceptron component to obtain the first forward propagation result. The multi-layer perceptron (MLP) is a neural network containing multiple hidden layers, which can learn complex nonlinear relationships between input vectors. In the full life cycle safety management of buses, the first standardized encoding vector contains information that has been internally focused and standardized, but this information may require further feature extraction and conversion to be better used for bus use specification evaluation. The multi-layer perceptron component processes the first standardized encoding vector through a series of linear transformations and nonlinear activation functions.

[0112] Assume that the multilayer perceptron has L hidden layers and the input of the lth layer is h l-1 (where h 0 =N1), then the output of the lth layer is , where W l is the weight matrix of the lth layer, b lis the bias vector of the lth layer, and f is a nonlinear activation function, such as the ReLU function (f(x)=max(0, x)). The ReLU function can introduce nonlinear factors, allowing the multilayer perceptron to learn more complex feature representations. After the forward propagation of the multilayer perceptron, the first forward propagation result P is finally obtained.

[0113] In step 40204, the first standardized encoding vector and the first forward propagation result corresponding to each training text embedding vector are standardized by the second cross-layer identity mapping component and the standardization component to obtain the training integrated text vector corresponding to each training text embedding vector. Similar to step 40202, the second cross-layer identity mapping component adds the first standardized encoding vector N1 to the first forward propagation result P to obtain R'=N1+P. This step again uses the cross-layer identity mapping to directly transfer the information of the previous layer to the current layer, which helps to alleviate the gradient vanishing and gradient exploding problems and retain more original information.

[0114] Then, the normalization component performs Layer Normalization on the added result, and after being processed by the second cross-layer identity mapping component and the normalization component, a training integrated text vector T=LN(R') corresponding to each training text embedding vector is obtained.

[0115] Through steps 40201-40204, the first training pre-coding vector is processed and optimized by the vector integration component. The internal attention component focuses on important information, the cross-layer identity mapping component and the standardization component improve the training stability of the model, and the multi-layer perceptron component learns the complex nonlinear relationship between the input vectors. The final training integrated text vector can more accurately reflect the usage characteristics of the bus in various archive record dimensions, providing stronger support for the subsequent generation of bus sample usage representation vectors and the evaluation of bus usage specifications.

[0116] As another implementation scheme for training the first text processing network and the second text processing network, the scheme includes: Step 1000: Obtain an archive record training text, a usage specification evaluation result supervision mark corresponding to a record entity in the archive record training text, and a relevance supervision mark of the archive record training text; the relevance supervision mark is used to characterize task execution information of a relevance task of the archive record training text; Step 2000: Load the archive record training text into the first text processing network to perform text vector embedding, and obtain the training text vector corresponding to the archive record training text; Step 3000: Load the training text vector into the second text processing network, perform full-connection classification mapping on the archive record training text, and obtain the usage specification evaluation result corresponding to the record entity in the archive record training text; Step 4000: Load the training text vector into the relevance task network, perform relevance task processing on the archive record training text, and obtain the relevance task processing result corresponding to the archive record training text; Step 5000: Determine a training error value based on the use specification evaluation result and the use specification evaluation result supervision mark, as well as the correlation task processing result and the correlation supervision mark; Step 6000: Based on the training error value, perform convergence training on the network parameters of the first text processing network and the second text processing network.

[0117] In step 1000, the archive record training text, the supervision mark of the use specification evaluation result corresponding to the record entity in the archive record training text, and the relevance supervision mark of the archive record training text are obtained. The archive record training text is a text collection about the full life cycle management of the public vehicle, covering information in multiple dimensions such as vehicle service files, vehicle files, driver files and violation files. The supervision mark of the use specification evaluation result clarifies the result of whether the use of the public vehicle corresponding to each archive record training text is standardized, such as "standard", "basic standard" and "non-standard". The relevance supervision mark is used to characterize the task execution information of the relevance task of the archive record training text. For example, in some scenarios, it is necessary to judge the relevance between different information in the archive record training text, such as whether the annual inspection information in the vehicle service file and the speeding violation information in the violation file are related. The relevance supervision mark will record this related information. Taking the official vehicle of the unit as an example, the archive record training text may contain detailed records of a certain vehicle maintenance, the driver's violation ticket information, etc. The supervision mark of the use specification evaluation result will indicate whether the use of the vehicle during this period is in compliance with the specification, and the relevance supervision mark may record whether there is a potential connection between the maintenance time and the violation time.

[0118] In step 2000, the archival record training text is loaded into the first text processing network for text vector embedding, and the training text vector corresponding to the archival record training text is obtained. The first text processing network can use a pre-trained language model, such as BERT, XLNet, etc. These models can learn the semantic information and contextual relationship of the text, and convert the input archival record training text into a low-dimensional vector representation. Taking BERT as an example, it encodes the input text through a multi-layer Transformer architecture. Assuming that the archival record training text is a document containing multiple sentences, BERT will convert the words in each sentence into word vectors, and then calculate the correlation between words through the self-attention mechanism, and finally output the training text vector of the entire document. Specifically, the input of BERT is obtained by adding word embedding, position embedding and segment embedding. After being processed by multiple Transformer blocks, the context representation of each word is obtained, and the output corresponding to the [CLS] tag is taken as the training text vector of the entire document.

[0119] In step 3000, the training text vector is loaded into the second text processing network, and the archive record training text is fully connected and classified and mapped to obtain the use specification evaluation result corresponding to the record entity in the archive record training text. The second text processing network can be a multi-layer perceptron (MLP), which consists of an input layer, a hidden layer, and an output layer. The input layer receives the training text vector output by the first text processing network, and the hidden layer processes the input through a series of linear transformations and nonlinear activation functions (such as ReLU functions) to extract higher-level features. The number of neurons in the output layer is the same as the number of categories of the use specification evaluation result, and the output is converted into a probability distribution through the Softmax function. Assuming that the use specification evaluation result is divided into three cases: "standard", "basic standard", and "non-standard", the output layer has 3 neurons. After being processed by the Softmax function, the probability distribution P=[p1, p2, p3] obtained represents the probability that the bus use corresponding to the archive record training text is "standard", "basic standard", and "non-standard". The category with the largest probability is selected as the final use specification evaluation result.

[0120] In step 4000, the training text vector is loaded into the relevance task network, and the archive record training text is subjected to relevance task processing to obtain the relevance task processing result corresponding to the archive record training text. The relevance task network can be a specially designed neural network for processing the relevance between different information in the archive record training text. For example, it can determine whether they are related by calculating the similarity between vectors of information of different dimensions. Assume that the vector corresponding to the information of the vehicle service archive dimension is , the vector corresponding to the information of the violation file dimension is , the correlation task network can use cosine similarity to calculate the similarity between them , according to the similarity value, it is judged whether the information of the two dimensions is related. The relevance task network processes the training text vector and outputs the relevance task processing result, such as related or not related.

[0121] In step 5000, the training error value is determined based on the usage specification evaluation result and the usage specification evaluation result supervision mark, as well as the correlation task processing result and the correlation supervision mark. For the usage specification evaluation task, the cross entropy loss function L1 is used to measure the difference between the predicted usage specification evaluation result and the actual usage specification evaluation result supervision mark. The formula of the cross entropy loss function refers to the description of the aforementioned step 50 and is not repeated here. For the correlation task, the mean square error loss function can be used to measure the difference between the correlation task processing result and the correlation supervision mark. The formula of the mean square error loss function is , where r j is the jth element of the relevance supervision tag, is the jth element of the result of the correlation task, and n is the number of samples. The final training error value L=L1+L2 combines the errors of the normative evaluation task and the correlation task.

[0122] In step 6000, based on the training error value, the network parameters of the first text processing network and the second text processing network are trained for convergence. The back propagation algorithm is used to calculate the gradient of the training error value to the network parameters (such as the weight matrix and the bias vector), and then the parameters are updated according to the gradient. The feasible optimization algorithms include stochastic gradient descent (SGD), Adagrad, Adadelta, Adam, etc. Through multiple iterations of training, the network parameters are continuously adjusted so that the training error value is gradually reduced until the convergence condition is reached (such as the error value is less than a certain threshold or the number of iterations reaches the upper limit). At this time, the performance of the network is optimal, and the use specifications of the bus can be accurately evaluated and the relevance task in the archival record training text can be processed. Through this training scheme, the use specification evaluation task and the relevance task are combined, so that the first text processing network and the second text processing network can learn more abundant information during the training process. The relevance task can help the network better understand the relationship between different information in the archival record training text, thereby improving the accuracy of the use specification evaluation. At the same time, through joint training, the network can share information and promote each other between the two tasks, so that the final model has better performance in the safety management of the whole life cycle of the bus.

[0123] In the above training scheme, a multi-task learning mechanism is adopted. When training the network, the archival record training text used for training the network and its corresponding supervision mark of the use specification evaluation result are obtained, and the relevance supervision mark of the archival record training text is also obtained; then, after the text vector is extracted from the archival record training text, the obtained training text vector is loaded into the second text processing network for evaluation and classification, and loaded into the relevance task network for associated task execution; then, after the training error values ​​obtained from the two tasks are fused, the network is trained for convergence. In this way, a multi-task collaborative training mechanism can be used to improve the semantic understanding effect of the network, thereby improving the accuracy of the standard evaluation results.

[0124] An embodiment of the present application provides a computer system, including a memory and a processor, wherein the memory stores a computer program that can be executed on the processor, and when the processor executes the program, some or all of the steps in the above method are implemented.

[0125] The embodiment of the present application provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, some or all of the steps in the above method are implemented. The computer-readable storage medium can be transient or non-transient.

[0126] An embodiment of the present application provides a computer program, including a computer-readable code. When the computer-readable code is run in a computer device, a processor in the computer device executes some or all of the steps for implementing the above method.

[0127] The embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and when the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. The computer program product can be implemented specifically by hardware, software or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium, and in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK) and the like.

[0128] It should be noted here that the description of the various embodiments above tends to emphasize the differences between the various embodiments, and the same or similar aspects can be referenced to each other. The description of the above device, storage medium, computer program and computer program product embodiments is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the embodiments of the device, storage medium, computer program and computer program product of this application, please refer to the description of the method embodiment of this application for understanding.

[0129] Figure 2 A hardware entity diagram of a computer system provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, the hardware entity of the computer system 1000 includes: a processor 1001 and a memory 1002, wherein the memory 1002 stores a computer program that can be run on the processor 1001, and the processor 1001 implements the steps in the method of any of the above embodiments when executing the program.

[0130] The memory 1002 stores computer programs that can be run on the processor. The memory 1002 is configured to store instructions and applications executable by the processor 1001. It can also cache data to be processed or processed by the processor 1001 and various modules in the computer system 1000 (for example, image data, audio data, voice communication data, and video communication data). This can be achieved through flash memory (FLASH) or random access memory (Random Access Memory, RAM).

[0131] When the processor 1001 executes the program, the steps of any one of the above-mentioned methods for managing the safety of the entire life cycle of a bus combined with big data are implemented. The processor 1001 controls the overall operation of the computer system 1000.

[0132] An embodiment of the present application provides a computer storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the steps of the bus life cycle safety management method combined with big data as in any of the above embodiments.

[0133] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For the technical details not disclosed in the storage medium and device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding. The above processor can be at least one of a target application integrated circuit (Application Specific Integrated Circuit, ASIC), a digital signal processor (Digital Signal Processor, DSP), a digital signal processing device (Digital Signal Processing Device, DSPD), a programmable logic device (Programmable Logic Device, PLD), a field programmable gate array (Field Programmable Gate Array, FPGA), a central processing unit (Central Processing Unit, CPU), a controller, a microcontroller, and a microprocessor. It can be understood that the electronic device that realizes the above processor function can also be other, and the embodiments of the present application are not specifically limited.

[0134] The above-mentioned computer storage medium / memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface memory, an optical disk, or a compact disc read-only memory (CD-ROM) and the like; it can also be various terminals including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0135] It should be understood that in various embodiments of the present application, the size of the sequence number of each step / process above does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiment of the present application. The above-mentioned sequence number of the embodiment of the present application is only for description and does not represent the advantages and disadvantages of the embodiment. It should be noted that, in this article, the term "include", "comprise" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements that are not clearly listed, or also includes elements inherent to such process, method, article or device. In the absence of more restrictions, the elements defined by the sentence "including one..." do not exclude the presence of other identical elements in the process, method, article or device including the element.

[0136] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0137] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, etc., various media that can store program codes.

[0138] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can essentially or in other words, the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0139] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.

Claims

1. A bus life cycle safety management method combined with big data, characterized in that: The method comprises: Obtaining the historical archive record text of the bus life cycle archive, wherein the bus life cycle archive is a record set for the full life cycle management of the target bus; Extracting archive text segments of multiple archive record dimensions from the historical archive record text, and mapping the extracted multiple archive text segments into multiple text embedding vectors; wherein the multiple archive record dimensions include a vehicle service archive dimension, a vehicle archive dimension, a driver archive dimension, and a violation archive dimension; Using the first vector coverage range as the scaling size and the second vector coverage range as the scaling step, performing a moving scaling process on each of the text embedding vectors to obtain a first text sub-vector set corresponding to each of the text embedding vectors; Performing vector integration processing on each of the first text sub-vector sets, and combining a plurality of integrated text vectors obtained by integration to obtain a bus usage representation vector of the bus life cycle archive; The use specification evaluation result of the bus life cycle file is determined by the bus use characterization vector.

2. The method according to claim 1, characterized in that The performing vector integration processing on each of the first text sub-vector sets, and combining the multiple integrated text vectors obtained by integration to obtain the bus usage representation vector of the bus life cycle file, includes: Determine a first text processing network corresponding to each of the first text sub-vector sets; Loading the first text sub-vector set corresponding to each of the text embedding vectors into the corresponding first text processing network for vector integration processing to obtain an integrated text vector corresponding to each of the text embedding vectors; Vector combination is performed on a plurality of the integrated text vectors to obtain a bus usage representation vector of the bus life cycle archive.

3. The method according to claim 2, characterized in that The step of combining the plurality of integrated text vectors to obtain the bus usage representation vector of the bus life cycle file includes: Determining an evaluation contribution coefficient of each of the integrated text vectors; Performing weighted adjustment on each of the integrated text vectors by using the evaluation contribution coefficient to obtain a plurality of adjusted text vectors; Performing vector combination on the multiple adjustment text vectors to obtain a bus usage representation vector of the bus life cycle file; The process of determining the evaluation contribution coefficient includes: Get the vehicle authorization level; According to the correlation information between each of the archive record dimensions and the vehicle weighting level, the evaluation contribution coefficient of each of the integrated text vectors is determined.

4. The method according to claim 1, characterized in that After extracting archive text segments of multiple archive record dimensions from the historical archive record text, and mapping the extracted multiple archive text segments into multiple text embedding vectors, the method further includes: Using the third vector coverage as the scaling size and the fourth vector coverage as the scaling step, performing a moving scaling process on each of the text embedding vectors to obtain a second text sub-vector set corresponding to each of the text embedding vectors; The performing vector integration processing on each of the first text sub-vector sets, and combining the multiple integrated text vectors obtained by integration to obtain the bus usage representation vector of the bus life cycle file, includes: Performing vector integration processing on each of the first text sub-vector sets, and combining the multiple first integrated sub-text vectors obtained by integration to obtain a first bus use sub-representation vector, and performing vector integration processing on each of the second text sub-vector sets, and combining the multiple second integrated sub-text vectors obtained by integration to obtain a second bus use sub-representation vector; The first bus usage sub-characterization vector and the second bus usage sub-characterization vector are combined to obtain a bus usage characterization vector of the bus life cycle file.

5. The method according to claim 2, characterized in that: The determining of the usage specification evaluation result of the bus life cycle file by using the bus usage characterization vector includes: Loading the bus usage representation vector into a second text processing network for full connection mapping to obtain a first evaluation probability distribution; An evaluation result of the usage specification of the bus life cycle file is determined according to the first evaluation probability distribution.

6. The method according to claim 5, characterized in that The first text processing networks and the second text processing networks are obtained by parallel training through bus life cycle archive samples. The first text processing networks and the second text processing networks are trained through the following process: Obtaining archive record training text, wherein the archive record training text includes full-cycle archive record training text of multiple bus training examples and a supervision mark of a usage specification evaluation result of each bus training example; Extracting archive training text segments of multiple archive record dimensions from the full-cycle archive record training text, and mapping the extracted multiple archive training text segments into multiple training text embedding vectors; Using the first vector coverage range as the scaling size and the second vector coverage range as the scaling stride, performing mobile scaling processing on each of the training text embedding vectors to obtain a first training text sub-vector set corresponding to each of the training text embedding vectors; Loading the first training text sub-vector set corresponding to each of the training text embedding vectors into the corresponding first text processing network for vector integration processing, and combining the training integrated text vectors corresponding to each of the training text embedding vectors obtained by integration to obtain a bus sample usage representation vector; The bus sample usage representation vector is loaded into the second text processing network for full connection mapping to obtain a second evaluation probability distribution, and a mapping error value is determined by the second evaluation probability distribution and the corresponding usage specification evaluation result supervision mark; Convergence training is performed on a plurality of network parameters of the first text processing network and the second text processing network using the mapping error value.

7. The method according to claim 6, characterized in that After extracting archive training text segments of multiple archive record dimensions from the full-cycle archive record training text, and mapping the extracted multiple archive training text segments into multiple training text embedding vectors, the method further includes: Using the third vector coverage as the scaling size and the fourth vector coverage as the scaling step, performing mobile scaling processing on each of the training text embedding vectors to obtain a second training text sub-vector set corresponding to each of the training text embedding vectors; The step of loading the first training text sub-vector set corresponding to each of the training text embedding vectors into the corresponding first text processing network for vector integration processing, and combining the training integrated text vectors corresponding to each of the training text embedding vectors obtained by integration to obtain a bus sample usage representation vector includes: Loading the first training text sub-vector set corresponding to each of the training text embedding vectors into the corresponding first text processing network for vector integration processing, and combining the first training integrated sub-text vectors corresponding to each of the training text embedding vectors obtained by integration to obtain a first training bus usage sub-representation vector; Loading the second training text sub-vector set corresponding to each of the training text embedding vectors into the corresponding first text processing network for vector integration processing, and combining the second training integrated sub-text vectors corresponding to each of the training text embedding vectors obtained by integration to obtain a second training bus usage sub-representation vector; Combining the first training bus usage sub-characterization vector with the second training bus usage sub-characterization vector to obtain a bus sample usage characterization vector of the bus training sample; Each of the first text processing networks includes a text precoding component and a vector integration component, and the first training text sub-vector set corresponding to each of the training text embedding vectors is loaded into the corresponding first text processing network for vector integration processing, and the training integrated text vectors corresponding to each of the training text embedding vectors obtained by integration are combined to obtain a bus sample usage representation vector, including: Performing text precoding on a first training text sub-vector set corresponding to each training text embedding vector by the text precoding component to obtain a first training precoding vector; Performing vector integration processing on the first training precoding vector corresponding to each of the training text embedding vectors through the vector integration component to obtain a training integrated text vector corresponding to each of the training text embedding vectors; The training integrated text vectors corresponding to each of the training text embedding vectors are combined to obtain a bus sample usage representation vector.

8. The method according to claim 7, characterized in that Each of the first text processing networks further includes a position embedding component, and the text precoding component performs text precoding on a first training text sub-vector set corresponding to each of the training text embedding vectors. After obtaining the first training precoding vector, the method further includes: Performing position embedding on the first training text sub-vector set corresponding to each of the training text embedding vectors through the position embedding component to obtain a first training text position sub-vector set; The step of performing vector integration processing on the first training precoding vector corresponding to each of the training text embedding vectors by the vector integration component to obtain a training integrated text vector corresponding to each of the training text embedding vectors includes: The vector integration component performs vector integration processing on the first training text position sub-vector set corresponding to each of the training text embedding vectors to obtain a training integrated text vector corresponding to each of the training text embedding vectors.

9. The method according to claim 8, characterized in that The step of performing position embedding on a first training text sub-vector set corresponding to each training text embedding vector by the position embedding component to obtain a first training text position sub-vector set includes: Determining the importance of the archive dimension corresponding to each of the training text sub-vector sets; The first training text position sub-vector set corresponding to each training text embedding vector is positionally embedded according to the archive dimension importance to obtain a first training text position sub-vector set.

10. The method according to claim 7, characterized in that The vector integration component includes an internal attention component, a first cross-layer identity mapping component and a normalization component, a multi-layer perceptron component, a second cross-layer identity mapping component and a normalization component. The vector integration component performs vector integration processing on the first training precoding vector corresponding to each training text embedding vector to obtain a training integrated text vector corresponding to each training text embedding vector, including: Performing internal weight focusing on the first training pre-coding vector corresponding to each of the training text embedding vectors through the internal attention component to obtain a first focused coding vector; Performing a standardization operation on the first focused encoding vector and the first training pre-encoding vector corresponding to each of the training text embedding vectors through the first cross-layer identity mapping component and the standardization component to obtain a first standardized encoding vector; Performing forward propagation on the first standardized encoding vector corresponding to each of the training text embedding vectors through the multi-layer perceptron component to obtain a first forward propagation result; The first standardized encoding vector and the first forward propagation result corresponding to each of the training text embedding vectors are standardized by the second cross-layer identity mapping component and the standardization component to obtain a training integrated text vector corresponding to each of the training text embedding vectors.

11. A computer system comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, wherein: When the processor executes the program, the steps in the method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Method and system for evaluating condition disorder of chronic disease patient

    CN116687410A

  • Malicious comment detection method based on Bert and Bi-LSTM

    CN116881449A

  • Automobile loan overdue day number prediction method fusing multiple types of features

    CN118864080A

  • Dangerous behavior identification and early warning method based on multi-modal analysis

    CN119360278A

  • Target personnel confidentiality awareness assessment method and system based on multivariate sentiment analysis

    CN119719366A