Intelligent vehicle moving method and system based on large model and storage medium

Through the intelligent car moving method based on large models, the problem of low accuracy in information extraction in the existing technology is solved, efficient identification of license plates and addresses is achieved, service efficiency and user experience are improved, and it is suitable for intelligent car moving systems.

CN120277189APending Publication Date: 2025-07-08KEXUN JIALIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510366790.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing technology has problems with expression diversity, difficulty in exhausting rules, high training costs of traditional BERT models and insufficient real-time performance in intelligent vehicle mobility services, resulting in low accuracy of information extraction, affecting service efficiency and user experience.

Method used

The intelligent car moving method based on big models is adopted to collect and preprocess user call data, establish and identify large models, identify license plate and address information, and use BP neural network or RBF neural network model for training and adjustment, and combine multimodal data sets for feature learning to enhance the fault tolerance and robustness of the model.

Benefits of technology

It realizes efficient and accurate identification of license plates and addresses in a short period of time, improves service efficiency and user satisfaction, reduces manual intervention, supports multiple concurrent interpretation of user needs, and improves the real-time and accuracy of transportation services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277189A_ABST
    Figure CN120277189A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent vehicle moving method and system based on a large model and a storage medium, relates to the technical field of intelligent vehicle moving, and solves the technical problem that the service efficiency and the user experience are affected due to the fact that an existing vehicle moving service system cannot achieve high information extraction accuracy within a short time. The method comprises the following steps: collecting call data of a user, and preprocessing; establishing an identification model and identifying promot of the large model according to the call data; the identification model identifies the information of the vehicle to be moved from the call data; according to the information of the to-be-moved vehicle, obtaining a vehicle owner number, feeding back to the user side, and completing the vehicle moving service; according to the method, the license plate and the address in man-machine conversation are accurately extracted by fusing the large model, and the time efficiency of vehicle moving appeal processing is accelerated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent vehicle relocation, and specifically relates to an intelligent vehicle relocation method, system, and storage medium based on a large model. Background Art

[0002] Intelligent vehicle relocation solutions involve intelligent robots providing services to meet the vehicle relocation demands of citizens, in forms such as phone hotlines, online text, etc. Existing technologies extract license plate numbers, addresses, etc. through keyword, rule, or traditional BERT model extraction based on the content described by citizens.

[0003] However, the existing technologies have the following significant drawbacks: 1. Problem of expression diversity: When citizens describe license plates, there are often various different expression forms, such as abbreviations, colloquial expressions, typos, etc. Keyword and rule extraction technologies are difficult to cover all possible expression forms, resulting in limited accuracy of information extraction; 2. Difficulty in exhaustive rules: Due to the diversity and complexity of license plate expressions, keyword and rule extraction technologies are difficult to exhaust all possible expression rules, leading to omissions or misjudgments in actual applications; 3. Limitations of the traditional BERT model: Although the BERT model performs well in natural language processing tasks, its training requires a large amount of labeled data and has a high training cost; 4. Real-time requirement: Vehicle relocation services usually require quick responses. Existing technologies often cannot achieve a high information extraction accuracy within a short time when processing citizens' vehicle relocation demands, affecting the service efficiency and user experience.

[0004] Therefore, the present invention provides an intelligent vehicle relocation method, system, and storage medium based on a large model. Summary of the Invention

[0005] The present invention aims to solve at least one of the technical problems existing in the prior art; for this purpose, the present invention proposes an intelligent vehicle relocation method, system, and storage medium based on a large model to solve the technical problem that the existing vehicle relocation service system cannot achieve a high information extraction accuracy within a short time, affecting the service efficiency and user experience.

[0006] To achieve the above object, the first aspect of the present invention provides an intelligent vehicle relocation method based on a large model, including the following steps:

[0007] Collect the call data of users and perform preprocessing; establish an identification model and a prompt for identifying the large model based on the call data;

[0008] The identification model identifies the information of the vehicle to be relocated from the call data; wherein, the information of the vehicle to be relocated includes the license plate and address of the vehicle to be relocated;

[0009] After obtaining the owner's number based on the information of the vehicle to be relocated, feedback it to the user terminal to complete the vehicle relocation service.

[0010] It should be noted that considering that the vehicle owner's number may be a foreign number, a call needs to be established according to the service specifications of the operator line provider.

[0011] Preferably, the call data includes text, voice text, and picture text; the prompts of the recognition large model include text type, voice type, and picture type.

[0012] Preferably, establishing the recognition model according to the call data includes:

[0013] The recognition model is constructed by an artificial intelligence model, and the construction process is as follows:

[0014] Obtain a number of data of vehicles to be moved and call tags from historical vehicle relocation data;

[0015] Among them, match the call tags with the corresponding call data;

[0016] Integrate the call data and call tags into several groups of data sets respectively; among them, the data set includes training data and test data;

[0017] Use the training data to train the artificial intelligence model; use the test data to test the trained artificial intelligence model, and adjust the artificial intelligence model according to the test results; finally obtain a recognition model with call tags as input and call data as output; among them, the artificial intelligence model is a BP neural network model or an RBF neural network model.

[0018] Preferably, the data set includes text and corresponding call tags, voice text and corresponding call tags, and pictures and corresponding call tags.

[0019] By integrating three types of heterogeneous data, namely text, voice text, and pictures, the model can simultaneously capture semantic information (text), voice features (tone, speech rate), and visual cues (scene images) to form complementary judgments. The data set contains multi-modal data corresponding to three types of tags. During training, the model is forced to learn cross-modal association features, enhancing the fault tolerance ability for noisy data, such as fuzzy voice and occluded pictures.

[0020] Preferably, the preprocessing process of the text includes the following steps:

[0021] Based on the statistical byte pair encoding algorithm, split the original text into sub-word units; use regular expression matching and stop word list to remove meaningless words; apply the lemmatization and stemming algorithms to map the vocabulary to the dictionary base form; convert the text sequence into a numerical vector through a pre-trained tokenizer, and pad or truncate the sequence to a preset length threshold.

[0022] Among them, the meaningless words include non-semantic characters and redundant stop words.

[0023] Preferably, the preprocessing process of the speech text includes the following steps:

[0024] Divide the continuous speech signal into short-time frames according to a preset time window; calculate the Mel-frequency cepstral coefficients or logarithmic Mel spectrograms of each frame of speech to generate a time-frequency feature matrix, and perform normalization on the feature matrix; convert the normalized features into a three-dimensional tensor; where the three-dimensional tensor is the time frame, frequency dimension, and number of channels.

[0025] Preferably, the preprocessing process of the picture text includes the following steps:

[0026] Scale the input image to the target resolution based on the bilinear interpolation algorithm; linearly map the pixel values from the original interval to a preset interval, and calculate the mean-standard deviation normalization by channel; perform channel replication on the grayscale image to generate pseudo RGB three-channel data; convert the image data into a four-dimensional tensor; where the four-dimensional tensor is the number of images, height, width, and number of channels.

[0027] Before inputting into the large model, the present invention preprocesses multi-modal data such as speech, text, and pictures, which can systematically improve the model performance and robustness. First, for noise reduction, voice activity detection, and speech rate normalization of speech text, the quality of the acoustic signal can be significantly enhanced; for post-correction of OCR and image enhancement technology of picture text, the text information under blurred or low-light conditions can be repaired, and the accuracy of visual feature extraction can be improved. In addition, the cross-modal time alignment and modal missing compensation mechanism can enhance the complementarity of multi-modal features; finally, the standardized input accelerates the model convergence speed and improves the robustness.

[0028] Preferably, the picture text generates enhanced samples by random rotation or brightness adjustment.

[0029] Preferably, a second aspect of the present invention provides an intelligent vehicle moving system based on a large model, including a configuration module, an identification module, and a completion module;

[0030] Configuration module: Collect the call data of the user and perform preprocessing; establish an identification model and a promot for the identification large model according to the call data;

[0031] Identification module: Used to identify the vehicle information to be moved from the call data based on the identification model; where the vehicle information to be moved includes the license plate and address of the vehicle to be moved;

[0032] Completion module: Used to obtain the owner's number according to the vehicle information to be moved and feedback it to the user terminal to complete the vehicle moving service.

[0033] Preferably, the third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement the method of the present invention.

[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0035] In the present invention, the large model technology is used to automatically and accurately identify and understand the vehicle moving type, extract the vehicle license plate number and the accident address in the vehicle moving service scenario. The large model can be called flexibly, and different task prompt word configurations are completed according to the business scenario. The large model extracts information from multiple vehicle moving service conversations in real-time and multi-concurrency, no longer relying on a large number of manual operators to answer citizens' vehicle moving calls or online requests. Each robot concurrently interprets the conversation intention and completes the question answering; the large model automatically extracts the elements required for the work order according to the conversation content, and there is no need for manual collation and entry into the system based on the recording or communication record; the large model extracts the summary of the user conversation to help the employees of the service provider quickly query the service records and classify and summarize the statistics; and the conversation interaction form includes multi-modal information such as text, voice, image, location, etc., realizing a comprehensive perception of the user's needs. By using the large model to complete the recognition of the vehicle moving type intention described by citizens, license plate and address extraction, the efficiency of vehicle moving services during peak traffic periods and holidays is greatly improved, and the satisfaction of citizens with traffic services is enhanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0037] Figure 1 It is a schematic flow chart of the method of the present invention;

[0038] Figure 2 It is a schematic structural diagram of the system of the present invention;

[0039] Figure 3 It is a schematic work flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0041] Please refer to Figure 1, an embodiment of the first aspect of the present invention provides an intelligent vehicle moving method based on a large model, including the following steps:

[0042] Obtain the call data of the user and preprocess the call data; wherein, the call data includes text, speech text, and picture text.

[0043] Preprocess the text, including the following processes:

[0044] Through the byte pair encoding algorithm based on statistics, split the original text into sub-word units; use regular expression matching and stop word list to remove meaningless words; apply lemmatization and stemming algorithms to map the vocabulary to the dictionary base form; convert the text sequence into a numerical vector through a pre-trained tokenizer, and pad or truncate the sequence to a preset length threshold.

[0045] It should be noted that stop words refer to words that frequently appear in the text but contribute little to understanding the content, such as "de" (of), "shi" (is), etc. In this specific example, "license plate number" and "address" are both key information and are not regarded as stop words. If there is a word like "de" in the sentence, it will be recognized and removed by regular expression matching and the stop word list.

[0046] Lemmatization attempts to change the word back to its dictionary form, while stemming simply removes the affixes for simplification purposes. "License plate number" and "address" are already in the base form, so they will not change in this step.

[0047] Pre-trained models usually have a vocabulary that can map the input text to the corresponding numerical representation. Assuming the vocabulary already contains "license plate number" and "address", then these two phrases will be converted into the corresponding vector representations.

[0048] Finally, adjust the lengths of all sequences uniformly. If the set maximum length is 10 tokens and the current sequence is less than 10 tokens, a specific padding token will be used; if it exceeds 10 tokens, it may be truncated from the front or the back.

[0049] Divide the continuous speech signal into short-time frames according to a preset time window; calculate the Mel frequency cepstral coefficients or logarithmic Mel spectrograms of each frame of speech to generate a time-frequency feature matrix, and perform normalization on the feature matrix; convert the normalized features into a three-dimensional tensor.

[0050] Among them, the speech signal is segmented into small segments (frames) for easy feature extraction. Generally, the time window length is set between 20 - 40 milliseconds, and in order to capture the changes in the speech signal, there will be a certain degree of overlap between these frames (usually 10 milliseconds).

[0051] For example, if there is a 5 - second long voice file, with a 25 - millisecond time window and a 10 - millisecond step, approximately 480 frames will be obtained.

[0052] Assume that the calculation of MFCCs is selected. Then each voice frame will correspond to an MFCC vector; for all frames, a two - dimensional matrix is constructed; where, one row represents the MFCC vector of one frame.

[0053] For example: Assume there are 480 frames, and each frame has 13 MFCC coefficients. Then a 480×13 matrix will be formed, and this matrix is normalized (such as standardized with a mean of 0 and a variance of 1).

[0054] The last step is to convert the normalized two - dimensional feature matrix into a three - dimensional tensor, usually an adjustment made to adapt to the input format of certain types of neural networks.

[0055] For example, an audio is segmented into 480 frames, and 13 MFCC features are extracted for each frame. Its tensor is (1, 480, 13).

[0056] It should be noted that the accuracy of license plate extraction is a key link in the car - moving service. Especially during the voice input process, user descriptions have colloquial characteristics. For example, "J" is pronounced as "gou", "Q" is pronounced as "quan", "R" is pronounced as "2". Therefore, it is necessary to confirm the easily confused sounds and the final license plate twice. For example, "Is the second digit the number 2 or the letter R?" and "Is your license plate Jing AJ2F13?"

[0057] Based on the bilinear interpolation algorithm, the input image is scaled to the target resolution; the pixel values are linearly mapped from the original interval to a preset interval, and the mean - standard deviation normalization is calculated for each channel; channel replication is performed on the grayscale image to generate pseudo - RGB three - channel data; the image data is converted into a four - dimensional tensor.

[0058] When performing model training, it is usually necessary to normalize the pixel values to a specific interval (such as [0, 1] or [-1, 1]) to accelerate the training process and improve the model performance. Assume that the original pixel value range is [0, 255], and it is linearly mapped to the [0, 1] interval. The mean and standard deviation are calculated separately for each color channel (red, green, blue), and these statistics are used to standardize the data of each channel. For example, if the mean of a certain channel is 128 and the standard deviation is 64, then all pixel values of this channel can be standardized by subtracting the mean and then dividing by the standard deviation.

[0059] Also, the deep learning model is designed to process color (RGB) images. If the input is a single-channel grayscale image, this single channel is copied three times to create a three-channel "pseudo" RGB image. For example, if the pixel value of an original grayscale image is 100, then the R, G, and B values at that position in the corresponding pseudo RGB image will all be set to 100.

[0060] Organize the processed image data into a format suitable for input into the deep learning model, that is, convert the data into a four-dimensional tensor with a shape of (batch_size, height, width, channels); where batch_size represents the number of images input into the model simultaneously, height and width are the height and width of the image (after scaling), and channels represent the number of color channels (3 for RGB images and 1 for grayscale images, but in this example, 3 channels have been generated through copying).

[0061] For example, if you have 10 pseudo RGB images with a size of 224×224, your input tensor will be a four-dimensional tensor with a shape of (10, 224, 224, 3).

[0062] Establish an identification model and its promot based on the call data; where promot includes text type, voice type, and picture type;

[0063] The identification model is constructed by an artificial intelligence model, and the construction process is as follows:

[0064] Obtain a number of data of vehicles to be moved and call tags from historical vehicle moving data;

[0065] Among them, match the call tags with the corresponding call data;

[0066] Integrate the call data and call tags into several groups of data sets respectively; where the data sets include training data and test data; the data sets include text texts and their corresponding call tags, voice texts and their corresponding call tags, and pictures and their corresponding call tags;

[0067] Use the training data to train the artificial intelligence model; use the test data to test the trained artificial intelligence model, and adjust the artificial intelligence model according to the test results; finally obtain an identification model with call tags as input and call data as output; where the artificial intelligence model is a BP neural network model or an RBF neural network model.

[0068] The identification model identifies the information of the vehicle to be moved from the call data; where the information of the vehicle to be moved includes the license plate and address of the vehicle to be moved.

[0069] According to the information of the vehicle to be moved, query the owner's number. After obtaining the owner's number, feedback it to the user terminal and establish a three-party call to complete the vehicle moving service.

[0070] Among them, the first party is the initiator of the call, such as a customer service staff, an application user, or a relevant person who needs to contact the vehicle owner; the second party is the system's automatic voice service; the third party is the vehicle owner himself / herself who is automatically called by the system.

[0071] Please refer to Figure 2 , The second aspect of the present invention provides an intelligent vehicle moving system based on a large model, including a configuration module, an identification module, and a completion module;

[0072] The configuration module collects data of the vehicle to be moved and establishes different types of identification models according to the types of the data of the vehicle to be moved;

[0073] The identification module identifies the information of the vehicle to be moved from the call data based on the identification model; among them, the information of the vehicle to be moved includes the license plate and address of the vehicle to be moved;

[0074] The completion module queries the owner's number of the vehicle to be moved according to the result returned by the large model, calls the vehicle owner, and automatically conducts the process of establishing a three-party call. After the three-party call ends, a service work order is automatically formed.

[0075] The third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement the method of the present invention.

[0076] Some of the data in the above formula is calculated by removing the dimension and taking its numerical value. The formula is obtained by software simulation of a large amount of collected data to get a formula closest to the actual situation; the preset parameters and preset thresholds in the formula are set by those skilled in the art according to the actual situation or obtained by simulating a large amount of data.

[0077] The working principle of the present invention: Please refer to Figure 3 , Configuring the large model promot1 is to establish an identification large model; the large model dialogue analysis 2 is to call the large model for analysis of the user's dialogue data. During the actual dialogue with the user, input the dialogue content and prompt into the identification large model, and the identification large model completes the process of task extraction, such as extracting the license plate number, address, etc.; extracting the license plate 3 is the process of the program parsing the license plate according to the result returned by the large model; extracting the address 4 is the process of the program parsing the address according to the result returned by the large model; calling the vehicle owner 5 is the process of the program automatically establishing a three-party call after querying the owner's number according to the license plate; work order completion 6 is that after the three-party call ends, a service work order is automatically formed and the process ends.

[0078] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.

Claims

1. An intelligent vehicle relocation method based on a large model, characterized in that, It includes a client, a cloud end, and a vehicle end; the following steps are included: Collect the call data of the user and perform preprocessing; establish an identification model and a promot for the large identification model based on the call data; The identification model identifies the vehicle information to be moved from the call data; among them, the vehicle information to be moved includes the license plate and address of the vehicle to be moved; After obtaining the owner's number based on the vehicle information to be moved, it is fed back to the user end to complete the vehicle moving service.

2. The intelligent vehicle relocation method based on a large model according to claim 1, characterized in that, The call data includes text, voice text, and picture text; the promot of the large identification model includes text type, voice type, and picture type.

3. The intelligent vehicle relocation method based on a large model according to claim 2, wherein, The establishment of the identification model based on the call data includes: The identification model is constructed by an artificial intelligence model, and the construction process is as follows: Obtain a number of vehicle data to be moved and call tags from historical vehicle moving data; Among them, match the call tags with the corresponding call data; Integrate the call data and call tags into several groups of data sets respectively; among them, the data set includes training data and test data; Use the training data to train the artificial intelligence model; use the test data to test the trained artificial intelligence model, and adjust the artificial intelligence model according to the test results; finally, obtain an identification model with call tags as input and call data as output; among them, the artificial intelligence model is a BP neural network model or an RBF neural network model.

4. The intelligent vehicle relocation method based on a large model according to claim 3, wherein The data set includes text and corresponding call tags, voice text and corresponding call tags, and pictures and corresponding call tags.

5. The intelligent vehicle relocation method based on a large model according to claim 2, characterized in that, The preprocessing process of the text includes the following steps: Based on the statistical byte pair encoding algorithm, split the original text into sub-word units; use regular expression matching and stop word list to remove meaningless words; apply lemmatization and stemming algorithms to map the vocabulary to the dictionary base form; convert the text sequence into a numerical vector through a pre-trained tokenizer, and pad or truncate the sequence to a preset length threshold.

6. The intelligent vehicle relocation method based on a large model according to claim 2, wherein, The preprocessing process of the voice text includes the following steps: Divide the continuous voice signal into short-time frames according to a preset time window; calculate the mel-frequency cepstral coefficients or log mel-spectrum of each frame of voice to generate a time-frequency feature matrix, and perform normalization on the feature matrix; convert the normalized features into a three-dimensional tensor.

7. The intelligent vehicle relocation method based on a large model according to claim 2, wherein The preprocessing process of the picture text includes the following steps: Scale the input image to the target resolution based on the bilinear interpolation algorithm; linearly map the pixel values from the original interval to a preset interval, and calculate the mean-standard deviation normalization by channel; perform channel replication on the grayscale image to generate pseudo-RGB three-channel data; convert the image data into a four-dimensional tensor.

8. The intelligent vehicle relocation method based on a large model according to claim 7, wherein, The picture text generates enhanced samples through random rotation or brightness adjustment.

9. An intelligent vehicle relocation system based on a large model, operating based on the intelligent vehicle relocation method based on a large model according to any one of claims 1-8, characterized in that, It includes a configuration module, an identification module, and a completion module; Configuration module: Collect the call data of the user and perform preprocessing; establish an identification model and a promot for the large identification model based on the call data; Identification module: Used to identify the vehicle information to be moved from the call data based on the identification model; among them, the vehicle information to be moved includes the license plate and address of the vehicle to be moved; Completion module: Used to obtain the owner's number based on the vehicle information to be moved, and then feed it back to the user end to complete the vehicle moving service.

10. A computer-readable storage medium, having a computer program stored thereon, characterized in that, The program is executed by a processor to implement the method according to any one of claims 1-8.