Carriage number identification method and device, electronic equipment and medium
By block encoding and AED model identification of train car number images, the missed recognition problem of traditional models when identifying continuous same characters is solved, high-precision car number recognition is achieved, and the intelligence and safety of railway transportation are improved.
Patent Information
- Application Number
- CN202510473346.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional AED models are prone to missed recognition problems when identifying train car numbers, especially when dealing with the same continuous characters, it is difficult to accurately identify car numbers.
A car number recognition method is adopted. By acquiring the target image, segmenting it into multiple image blocks, and rotating encoding and repetitive intensity weight are performed on each image block, and identification is carried out in combination with the encoder and decoder models (main path and auxiliary path) to ensure accurate identification of the car number.
This method can effectively identify the carriage number, reduce missed errors, improve identification accuracy, save human resources, and enhance the safety and intelligence level of railway transportation.
Smart Images

Figure CN120014652A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of carriage number recognition, and more specifically, to a method, device, electronic equipment and medium for recognizing a carriage number. Background Art
[0002] Currently, optical character recognition (OCR) has shifted from traditional multi-module cascade to end-to-end modeling. The latter can better cope with the impact of environmental changes because it reduces the intermediate links. However, when dealing with fixed-length information such as train car numbers, where there may be consecutive identical characters, traditional AED models are prone to missed recognition problems. Summary of the invention
[0003] The purpose of the embodiments of the present application is to provide a method, device, electronic device and medium for identifying a car number, so as to solve the above-mentioned problems existing in the prior art and to identify the accurate car number.
[0004] In a first aspect, a method for identifying a carriage number is provided, and the method may include: Obtain a target image containing the carriage number to be identified; Processing the target image to obtain a plurality of image blocks and image data and position data of each image block; For any image block, encoding the image data and corresponding position data of the image block to obtain a position code of the image block; The image data and the corresponding position code in each image block are input into a trained car number recognition model to obtain the target car number in the target image; the car number recognition model includes an encoder and a decoder consisting of a main path for determining a character sequence and an auxiliary path for determining the expected number of repetitions of each character in the character sequence.
[0005] In a possible implementation, the position encoding includes rotation encoding and repetition intensity weighting.
[0006] In a possible implementation, the process of determining the repetition intensity weight includes: For any image block, based on the image data of the image block and the corresponding position data, determine the image data of the adjacent image block; Combining the image data of the image block with the image data of any adjacent image block to obtain a plurality of image data pairs; For any image data pair, similarity calculation is performed on two image data in the image data pair to obtain cosine similarity; The softmax function is used to process each cosine similarity to obtain the repetition intensity weight.
[0007] In a possible implementation, the image data and the corresponding position code in each image block are input into a trained car number recognition model to obtain the target car number in the target image, including: Inputting the image data and the corresponding position code in each image block into the encoder for processing to obtain multiple vector sequences; Inputting multiple vector sequences into the main path for processing to obtain a character sequence; the character sequence includes multiple characters; Inputting multiple vector sequences into the auxiliary path for processing to obtain the expected number of repetitions of each character; The target car number is obtained by fusing the multiple characters in the character sequence with the expected number of repetitions of each character.
[0008] In one possible implementation, the encoder includes a coding embedding layer and multiple first processing layers; the first processing layer includes a first multi-head attention module, a first normalization module, a first bit-by-bit feedforward network module and a second normalization module connected in sequence.
[0009] In one possible implementation, the main path includes a decoding embedding layer, multiple second processing layers and a fully connected layer connected in sequence; the second processing layer includes a second multi-head attention module, a third normalization module, a second bit-by-bit feedforward network module and a fourth normalization module connected in sequence.
[0010] In a possible implementation, the auxiliary path is a 5x3 convolutional layer structure.
[0011] In a second aspect, a carriage number recognition device is provided, which may include: An acquisition unit, used for acquiring a target image containing a carriage number to be identified; A processing unit, used for processing the target image to obtain a plurality of image blocks and image data and position data of each image block; And, for any image block, encoding the image data and corresponding position data of the image block to obtain the position code of the image block; And, the image data and the corresponding position code in each image block are input into a trained car number recognition model to obtain the target car number in the target image; the car number recognition model includes an encoder and a decoder composed of a main path for determining a character sequence and an auxiliary path for determining the expected number of repetitions of each character in the character sequence.
[0012] In a third aspect, an electronic device is provided, the electronic device comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory, used to store computer programs; The processor is used to implement any method step described in the first aspect when executing the program stored in the memory.
[0013] In a fourth aspect, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, any method step described in the first aspect is implemented.
[0014] The present application provides a method for identifying a car number, including: obtaining a target image containing a car number to be identified; processing the target image to obtain multiple image blocks and image data and position data of each image block; for any image block, encoding the image data and corresponding position data of the image block to obtain the position code of the image block; inputting the image data and corresponding position code in each image block into a trained car number recognition model to obtain the target car number in the target image; the present application can accurately identify the car number based on the car number recognition model, without the need for additional manual inspection, saving a lot of time and human resources. Accurate car number recognition helps prevent safety hazards caused by information errors and provides strong support for the intelligent upgrade of the railway transportation industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0016] Figure 1 A system architecture diagram of a method for identifying a carriage number provided in an embodiment of the present application; Figure 2 A flow chart of a method for identifying a carriage number provided in an embodiment of the present application; Figure 3 An architecture diagram of the Vision Transformer provided in an embodiment of the present application; Figure 4 The architecture diagram of the car number recognition model provided in the embodiment of the present application; Figure 5 A schematic diagram of the structure of a carriage number recognition device provided in an embodiment of the present application; Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0017] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0018] The method for identifying a carriage number provided in the embodiment of the present application can be applied in Figure 1 In the system architecture shown in Figure 1 As shown, the system may include: a server and a terminal. The server may be a physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal may be a user equipment (UE) such as a mobile phone, smart phone, laptop, digital broadcast receiver, personal digital assistant (PDA), tablet computer (PAD), handheld device, vehicle-mounted device, wearable device, computing device or other processing device connected to a wireless modem, mobile station (MS), mobile terminal (MobileTerminal), etc. The terminal and the server may be directly or indirectly connected via wired or wireless communication, which is not limited in this application.
[0019] The terminal is used to collect a target image including the carriage number to be identified and send the target image to the server; The server is used to obtain a target image containing a car number to be identified so as to execute a car number identification method provided in the present application.
[0020] End-to-end modeling methods have become the mainstream in optical character recognition modeling. Compared with traditional optical character recognition, end-to-end methods involve fewer modules, reduce module cascade errors, and not only have high recognition accuracy, but also have good robustness to changes in the external environment. Optical character recognition is a process of converting images to text. In recent years, the model of the encoder-decoder structure based on the self-attention mechanism (abbreviated as AED / Attention-based Encoder Decoder model) has become an increasingly hot research topic. This model not only dominates the field of natural language processing, but also appears more and more in the field of vision.
[0021] The basic approach of the AED model is to divide the input image into blocks, and convert each small block of the image into a vector through linear projection. The dimension of each element in the vector is D. This vector can be regarded as the initial token of the image. Combined with the encoding of the image block, the image is finally represented as a sequence structure like text. Due to different formats such as fonts and font sizes, the feature length extracted by the model is often different from the actual character sequence length, so modeling needs to consider sequence alignment.
[0022] Since the train carriage numbers are fixed 7-digit Arabic numerals, the fonts painted on different carriages are also different. Experiments have found that if there are consecutive identical characters in the carriage number text, one of the consecutive characters may be missed. For example, 1466626 is recognized as 146626, and 2875523 is recognized as 287523.
[0023] Existing technology defects: 1) When aligning sequences, the AED model assigns fuzzy attention weights to consecutive identical characters (such as consecutive 6s in 1466626), resulting in missed recognition or incorrect prediction of the number of repetitions.
[0024] 2) Traditional position encoding cannot distinguish the physical position differences of consecutive identical characters, and it is difficult for the model to perceive the number of repeated characters.
[0025] In summary, the fonts used to spray train carriage numbers are diverse, and there is a high probability of consecutive identical characters (such as the XXX555X format). Missing recognition will lead to failure in carriage number verification, and the recognition process requires manual intervention, affecting railway transportation efficiency.
[0026] Therefore, the present application provides a method for identifying a car number to solve the above-mentioned problems existing in the prior art and to identify the accurate car number.
[0027] The preferred embodiments of the present application are described below in conjunction with the drawings in the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. In addition, the embodiments and features in the embodiments of the present application may be combined with each other if there is no conflict.
[0028] Figure 2 The following is a flow chart of a method for identifying a carriage number provided in an embodiment of the present application. Figure 2 As shown, the method may include: Step S210: Acquire a target image containing the carriage number to be identified.
[0029] Specifically, first, a target image containing the carriage number to be identified needs to be obtained in some way. This can be done by actually taking a photo or extracting it from an existing database.
[0030] In order to ensure that the target image is clear, the lighting is moderate, and the carriage number is not blocked or damaged, multiple pictures can be taken from different angles and then merged into a more comprehensive target image using image stitching technology.
[0031] Afterwards, the target image may be preprocessed, which may include: The text area in the target image can be identified to frame the area, and irrelevant information other than the car number (such as background, other vehicles, etc.) can be removed based on the text area. In other words, the target image can be cropped based on the text area to retain only the part containing the car number.
[0032] Step S220: Process the target image to obtain a plurality of image blocks and image data and position data of each image block.
[0033] Combination Figure 3 As shown, specifically, the deep learning model based on the Vision Transformer architecture divides the input target image into multiple image blocks of fixed sizes.
[0034] The image data corresponding to each image patch is flattened and converted into a feature vector of fixed dimension.
[0035] Add position data to each feature vector.
[0036] Add a class token to the image sequence composed of multiple feature vectors.
[0037] Step S230: For any image block, perform encoding processing on the image data and corresponding position data of the image block to obtain a position code of the image block.
[0038] Specifically, position encoding It can be expressed as:
[0039] in, For rotary encoding, is the repetition intensity weight, is a learnable parameter; are the coordinates of the corresponding image block in the target image.
[0040] Furthermore, A. For the position data (x, y) of any image block, the rotation code of the image block can be calculated by the following formula:
[0041]
[0042] Among them, i is the dimension index (starting from 0), d modelis the total dimension of the rotation encoding vector.
[0043] Since each image block position data (x, y) will produce two values (one from the sine function and the other from the cosine function), for any position data (x, y), you will get a length d model This vector is the rotation code of the image block.
[0044] B. Assume that the target image is evenly divided into 81 image blocks of size 16x16 (the target image size can be 144x144 pixels, 81=9×9, and each small block is 16x16). Then, each 16x16 image block is converted into a 512-dimensional feature vector through an embedding layer (here a convolutional layer is used). This means that the entire image is represented as a 9x9x512 tensor, where 9x9 represents the position grid of the image block, and 512 is the feature dimension of the image block corresponding to each position.
[0045] For any image block, find all its directly adjacent neighboring image blocks (up, down, left, right, or more, depending on the specific application scenario) based on the position data (such as coordinates in the grid) and image data (i.e., pixel values) of the image block.
[0046] The data of the image block is paired with the image data of each found adjacent image block to form a plurality of image data pairs, each pair including the image data of one image block and the image data of a corresponding adjacent image block.
[0047] For each image data pair, the cosine similarity between the two image data is calculated. Cosine similarity is an indicator that measures the cosine value of the angle between two vector directions, which reflects the similarity between the two vectors. Specifically, it can be understood as follows: for each image block feature vector F at position i i , calculate the feature vector F of the image block and its adjacent image block at position j j The cosine similarity between them. Cosine similarity measures the similarity between the directions of two non-zero vectors. The formula is as follows:
[0048] Among them, F i ⋅F j represents the vector dot product, which is the inner product of two vectors; and |||F i ∣∣ and ∣∣F j ∣∣ respectively represent vector F i and F j The denominator is used to standardize the inner product result to ensure that the similarity value is in the range of [-1, 1].
[0049] After that, the softmax function is used to normalize multiple cosine similarities to obtain a probability distribution, which reflects the relative strength of the similarity between image blocks at each position, that is, the repetition strength weight is obtained. The softmax function is defined as follows:
[0050] Among them, S(j) represents the normalized similarity weight of the jth position relative to position i, and e is the base of the natural logarithm. The result of this step is a matrix corresponding to the original 9x9 grid. Each element in the matrix represents the similarity weight of the image block at the corresponding position. The larger the value, the more similar the image block at this position is to the image block at the reference position i.
[0051] The repetition strength weight represents the relative similarity or “repetition strength” between different image patches.
[0052] In summary, for any image block, the rotation code and the repetition intensity weight corresponding to the image block are added together to obtain the position code of the image block.
[0053] Step S240: input the image data and the corresponding position code in each image block into the trained car number recognition model to obtain the target car number in the target image.
[0054] Combination Figure 4 As shown, the car number recognition model includes an encoder and a decoder. The decoder includes: a main path and an auxiliary path; the main path is used to determine the character sequence ( Figure 4 Only the primary path is shown in ); the secondary path is used to determine the expected number of repetitions of each character in the character sequence.
[0055] Step S240 may specifically include: Step S240-1: Input the image data and corresponding position code of each image block into the encoder for processing to obtain a vector sequence; each image data can be converted into a vector representation suitable for model processing. The image features are extracted through a convolutional neural network (CNN), and each image data is converted into a feature map sequence, and then further converted into a vector sequence suitable for self-attention mechanism processing. Among them, the encoder includes a coding embedding layer and multiple first processing layers; the first processing layer includes a first multi-head attention module, a first normalization module, a first bit-by-bit feedforward network module and a second normalization module connected in sequence.
[0056] Furthermore, Multi-Head Attention executes multiple self-attention calculation heads in parallel, each head uses a different linear transformation matrix , and then concatenate the outputs of multiple heads and undergo a linear transformation to obtain the final output. This method allows the model to capture information from different subspaces and enhance the model's representation capabilities. They represent the query, key, and value transformation matrices of the hth attention head, respectively. Here, h represents different attention calculation heads (h=1,...,H), where H is the total number of attention heads.
[0057] Step S240-2: input the vector sequence into the main path for processing to obtain a character sequence; the character sequence includes a plurality of characters; The main path includes a decoder embedding layer, multiple second processing layers, and a fully connected layer connected in sequence; the second processing layer includes a second multi-head attention module, a third normalization module, a second bit-by-bit feedforward network module, and a fourth normalization module connected in sequence. In this stage, the length of the output character sequence is forced to be 7, and non-numeric characters and super-long sequence generation are shielded by dynamic masks.
[0058] Masked multi-head attention and the fifth normalization module are also configured in the main path.
[0059] During training, the main path in the decoder is looped, where each time step , taking the hidden state of the previous time step (which may be a zero vector or a specific initialization vector at the beginning) and the current self-attention output as input, and predicting the probability distribution of the next character through a fully connected layer. The predicted score is converted into a probability using the Softmax function, that is, ,in, It is The characters generated by the time step, is a sequence of characters, It is The hidden state at the time step, and Is the corresponding character The learnable weights and biases.
[0060] Then, based on the predicted probability distribution, strategies such as greedy search (selecting the character with the highest probability) or beam search (retaining multiple candidate characters with higher probabilities) can be used to determine the actual generated character, which is used as the input of the next time step to continue generating the next character until the preset end condition is reached (such as generating a specific end tag or reaching the maximum length).
[0061] Step S240 - 3 : Input the vector sequence into the auxiliary path for processing to obtain the expected number of repetitions of each character.
[0062] For example: [0, 0.3, 0.8, 1.0] corresponds to the repeating trend of "6, 6, 6".
[0063] The auxiliary path shares the input data (decoder output data) with the main path, and also preprocesses and extracts features from the input data to convert it into a vector representation. For example, if the input is text, a vector sequence is obtained through word embedding; if it is an image, a feature map sequence is obtained through CNN. In order for the car number recognition model to capture information related to repetition, some special designs can be adopted in the feature extraction stage. For example, for text, you can consider using PositionEncoding to enhance the model's perception of character positions, because repeated characters often have certain patterns in position. For this application, a 5x3 convolution kernel or convolution layer structure can be designed to make it more sensitive to repeated local patterns.
[0064] In this method, an auxiliary path is used in parallel with the main path to predict the expected number of repetitions of each character. The auxiliary path can be based on architectures such as multi-layer perceptron (MLP) or convolutional neural network (CNN, suitable for processing image features). Taking MLP as an example, the input feature vector is transformed through multiple fully connected layers and nonlinear activation functions (such as ReLU), and finally the predicted expected number of repetitions is output through a fully connected layer. Assume that the input feature vector is , after a series of fully connected layers And the activation function Then, we get the predicted value ,in, is the prediction of the expected number of repetitions of the ith character. This expression describes how to predict the ith character through a series of fully connected layers and activation functions. i The expected number of repetitions of characters. First, input the feature vector f i Through the first fully connected layer W1 and through the activation function σ Processing, and then the results are passed to the subsequent fully connected layers and activation functions in sequence until the last layer W m Final output During the training process, by defining an appropriate loss function (such as mean squared error loss for continuous predictions or cross entropy loss if the number of repetitions is discretized), the model learns to accurately predict the expected number of repetitions.
[0065] Step S240-4: Merge multiple characters in the character sequence with the expected number of repetitions of each character to obtain a target car number.
[0066] Specifically, a cross-attention mechanism is used for the interaction between the auxiliary path and the main path. In the cross-attention, the output of the auxiliary path (features related to the predicted number of repetitions) is used as the key and value, and the features of the main path are used as the query. The specific calculation process is similar to the standard self-attention. For the query vector of the main path The self-attention output from the main path) and the key vector of the auxiliary path , value vector (the feature vector associated with the predicted number of repetitions), calculate the attention score matrix , where i and j are mainly used to identify the specific locations of the query vector and key vector, d k is the key vector dimension; then we get the cross attention output The cross attention output The information related to the number of repetitions is carried and fused with the self-attention output of the main path. By concatenating and fusing and then performing a linear transformation, the main path can focus on the repeated area when generating characters, thereby improving the accuracy and rationality of the generated character sequence, especially for cases containing repeated characters.
[0067] This method can be applied to railway freight systems to automatically identify carriage numbers and improve marshaling efficiency. It can also be applied to customs inspections to quickly verify cross-border train numbers and reduce manual verification.
[0068] Before executing step S240, the process of training the pre-trained car number recognition model may include: Generate a training set of carriage numbers: simulate carriage number images under different fonts, lighting, and contamination conditions, and explicitly increase the proportion of consecutive identical character samples (such as: XX5555X).
[0069] Generate carriage number training labels: mark the number of repetitions of each character (for example, 1466626[1,4,6,6,6,2,6], the number of repetitions is: [1,1,3,1,1]).
[0070] During the training process, the preset loss function is used for joint optimization. The preset loss function is as follows:
[0071] in, is the weight coefficient used to balance the contribution of the two losses, For cross entropy, a standard cross entropy loss is used to measure the difference between the generated characters and the true labels. is the mean square error, and the mean square error (MSE) loss is used to predict the difference between the number of repetitions and the true value.
[0072]
[0073] in, is the length of the character sequence, For the character table, is the time step The one-hot encoding of the real character, The main path at time step Predicted Characters probability.
[0074]
[0075] in, For the auxiliary path at time step Predict the number of repetitions, is the time step The actual number of repetitions.
[0076] Furthermore, the trained car number recognition model is verified: A total of 678 images containing carriage numbers to be identified were collected, of which 145 images contained carriage numbers with continuous characters. The method proposed in this application has a 10.1% lower continuous character miss rate than the traditional AED model, and the overall carriage number recognition rate has increased by 9.9%. The specific data is as follows:
[0077] The present application provides a method for identifying a car number, including: obtaining a target image containing a car number to be identified; processing the target image to obtain multiple image blocks and image data and position data of each image block; for any image block, encoding the image data and corresponding position data of the image block to obtain the position code of the image block; inputting the image data and corresponding position code in each image block into a trained car number recognition model to obtain the target car number in the target image; the present application can accurately identify the car number based on the car number recognition model, without the need for additional manual inspection, saving a lot of time and human resources. Accurate car number recognition helps prevent safety hazards caused by information errors and provides strong support for the intelligent upgrade of the railway transportation industry.
[0078] Corresponding to the above method, the embodiment of the present application also provides a car number recognition device, such as Figure 5 As shown, the device comprises: An acquisition unit 510 is used to acquire a target image including a carriage number to be identified; A processing unit 520 is used to process the target image to obtain a plurality of image blocks and image data and position data of each image block; And, for any image block, encoding the image data and corresponding position data of the image block to obtain the position code of the image block; And, the image data and the corresponding position code in each image block are input into a trained car number recognition model to obtain the target car number in the target image; the car number recognition model includes an encoder and a decoder composed of a main path for determining a character sequence and an auxiliary path for determining the expected number of repetitions of each character in the character sequence.
[0079] The functions of each functional unit of a car number identification device provided in the above embodiment of the present application can be achieved through the above method steps. Therefore, the specific working process and beneficial effects of each unit in a car number identification device provided in the embodiment of the present application will not be repeated here.
[0080] The present application also provides an electronic device, such as Figure 6 As shown, it includes a processor 610 , a communication interface 620 , a memory 630 and a communication bus 640 , wherein the processor 610 , the communication interface 620 , and the memory 630 communicate with each other via the communication bus 640 .
[0081] Memory 630, for storing computer programs; The processor 610 is used to implement the following steps when executing the program stored in the memory 630: Obtain a target image containing the carriage number to be identified; Processing the target image to obtain a plurality of image blocks and image data and position data of each image block; For any image block, encoding the image data and corresponding position data of the image block to obtain a position code of the image block; The image data and the corresponding position code in each image block are input into a trained car number recognition model to obtain the target car number in the target image; the car number recognition model includes an encoder and a decoder consisting of a main path for determining a character sequence and an auxiliary path for determining the expected number of repetitions of each character in the character sequence.
[0082] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0083] The communication interface is used for communication between the above electronic device and other devices.
[0084] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0085] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0086] The implementation methods and beneficial effects of the components of the electronic device in the above embodiments to solve the problems can be seen in Figure 2 The various steps in the illustrated embodiment are implemented, therefore, the specific working process and beneficial effects of the electronic device provided by the embodiment of the present application are not repeated here.
[0087] In another embodiment provided in the present application, a computer-readable storage medium is also provided, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer executes a method for identifying a car number described in any of the above embodiments.
[0088] In another embodiment provided in the present application, a computer program product comprising instructions is also provided, which, when executed on a computer, enables the computer to execute a method for identifying a car number described in any one of the above embodiments.
[0089] Those skilled in the art will appreciate that the embodiments in the present application can be provided as methods, systems, or computer program products. Therefore, the embodiments in the present application can be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments in the present application can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0090] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0091] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0092] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0093] Unless otherwise defined, the technical terms or scientific terms used in this application should be understood by people with ordinary skills in the field to which the present invention belongs. "First", "second" and similar words used in this application do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect", "couple" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0094] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they are aware of the basic creative concepts. Therefore, the present application embodiments are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application embodiments.
[0095] Obviously, those skilled in the art can make various changes and modifications to the embodiments in the present application without departing from the spirit and scope of the embodiments in the present application. Thus, if these modifications and variations of the embodiments in the present application are within the scope of the embodiments in the present application and their equivalents, the embodiments in the present application are also intended to include these modifications and variations.
Claims
1. A method for identifying a carriage number, characterized in that: The method comprises: Obtain a target image containing the carriage number to be identified; Processing the target image to obtain a plurality of image blocks and image data and position data of each image block; For any image block, encoding the image data and corresponding position data of the image block to obtain a position code of the image block; The image data and the corresponding position code in each image block are input into a trained car number recognition model to obtain the target car number in the target image; the car number recognition model includes an encoder and a decoder consisting of a main path for determining a character sequence and an auxiliary path for determining the expected number of repetitions of each character in the character sequence.
2. The method according to claim 1, characterized in that The position encoding includes rotation encoding and repetition intensity weighting.
3. The method according to claim 2, characterized in that The process of determining the repetition intensity weight includes: For any image block, based on the image data of the image block and the corresponding position data, determine the image data of the adjacent image block; Combining the image data of the image block with the image data of any adjacent image block to obtain a plurality of image data pairs; For any image data pair, similarity calculation is performed on two image data in the image data pair to obtain cosine similarity; The softmax function is used to process each cosine similarity to obtain the repetition intensity weight.
4. The method according to claim 1, characterized in that The image data and the corresponding position code in each image block are input into the trained car number recognition model to obtain the target car number in the target image, including: Inputting the image data and the corresponding position code in each image block into the encoder for processing to obtain multiple vector sequences; Inputting multiple vector sequences into the main path for processing to obtain a character sequence; the character sequence includes multiple characters; Inputting multiple vector sequences into the auxiliary path for processing to obtain the expected number of repetitions of each character; The target car number is obtained by fusing the multiple characters in the character sequence with the expected number of repetitions of each character.
5. The method according to claim 1, characterized in that The encoder includes a coding embedding layer and multiple first processing layers; the first processing layer includes a first multi-head attention module, a first normalization module, a first bit-by-bit feedforward network module and a second normalization module connected in sequence.
6. The method according to claim 1, characterized in that The main path includes a decoding embedding layer, multiple second processing layers and a fully connected layer connected in sequence; the second processing layer includes a second multi-head attention module, a third normalization module, a second bit-by-bit feedforward network module and a fourth normalization module connected in sequence.
7. The method according to claim 1, characterized in that The auxiliary path is a 5x3 convolutional layer structure.
8. A carriage number recognition device, characterized in that: The device comprises: An acquisition unit, used for acquiring a target image containing a carriage number to be identified; A processing unit, used for processing the target image to obtain a plurality of image blocks and image data and position data of each image block; And, for any image block, encoding the image data and corresponding position data of the image block to obtain the position code of the image block; And, the image data and the corresponding position code in each image block are input into a trained car number recognition model to obtain the target car number in the target image; the car number recognition model includes an encoder and a decoder composed of a main path for determining a character sequence and an auxiliary path for determining the expected number of repetitions of each character in the character sequence.
9. An electronic device, characterized in that: The electronic device comprises a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory, used to store computer programs; A processor, for implementing the method steps described in any one of claims 1 to 7 when executing a program stored in a memory.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Power transformation equipment overheating early warning method and system based on multi-modal fusion
CN117372845A
Train carriage information intelligent identification method and system
CN118799852A
Model training and image recognition methods and apparatuses, device, storage medium and computer program product
WO2023142551A1
KR20230075340A