Method and device of online car-hailing voice assistant system based on deep learning

The deep learning-based net car voice assistant system addresses inefficiencies in traditional systems by enhancing voice recognition and understanding through self-encoder models, ensuring faster and more accurate order processing.

CN120319239APending Publication Date: 2025-07-15BEIJING BAIJU YIXING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510525473.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Traditional voice assistant systems are difficult to meet the real-time response needs in terms of processing speed and recognition accuracy, and cannot effectively understand the user's context and emotional information, resulting in poor user experience.

Method used

Deep learning technology is adopted to feature engineering the online ride-hailing order data through the autoencoder model, generate training sets, test sets and verification sets, and conduct multiple rounds of training until the loss function converges, generate an autoencoder model, and input real-time data into the model for speech recognition and understanding.

Benefits of technology

It improves the accuracy and response speed of voice recognition, can better understand users' voice commands, provide convenient and efficient humanized services, reduce misunderstandings and error rates, and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120319239A_ABST
    Figure CN120319239A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a device of an online car-hailing voice assistant system based on deep learning, and particularly relates to the technical field of artificial intelligence. Online car-hailing order data is randomly selected from an online database, including recording of a user voice instruction and corresponding text conversion, and feature engineering processing is performed on the collected voice instruction data; the method comprises the following steps: generating a training set, a test set and a verification set, loading the training set and the test set, carrying out multiple rounds of training until a loss function converges, generating an auto-encoder model, reading verification set data, inputting the verification set data into the auto-encoder model, verifying the effect of the model on the verification set, and analyzing an output result. And inputting the data after real-time processing into a trained auto-encoder model to obtain a voice recognition and understanding result based on each order, and automatically processing the orders according to an output result of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and more specifically, to a method and device for a voice assistant system for online car-hailing based on deep learning. Background Art

[0002] With the rapid development of the digital age, online car-hailing services have gradually become the first choice for people's travel. Users' expectations for travel experience are constantly increasing, especially in the use of voice assistant systems. More and more users hope to complete functions such as hailing a car, payment, and route query through simple voice commands. However, traditional voice assistant systems face many challenges and need the introduction of innovative technologies.

[0003] Traditional voice assistant systems mainly rely on rule matching and keyword recognition. Although this method can meet basic functional requirements to a certain extent, there are obvious drawbacks in actual applications. First of all, the problem of insufficient timeliness greatly reduces the user experience. Since traditional voice recognition technology often relies on the gradual analysis of voice signals, its processing speed and recognition accuracy are difficult to meet the standards of real-time response. For users in urgent need of hailing a car, this undoubtedly increases the waiting time for travel and reduces service satisfaction.

[0004] Secondly, the problem of insufficient learning depth makes the system perform limited in understanding user intentions. Language is not only composed of simple words, but also contains rich context and emotional information. Traditional systems often cannot effectively understand the context relationship, resulting in misunderstandings or inability to respond to the real needs of users. For example, users may use different expressions to ask the same question, and traditional systems are difficult to accurately identify this diversity, causing communication barriers between users and the system.

[0005] In view of the above problems, a voice assistant system for online car-hailing based on deep learning has emerged as the times require. Through the training of large-scale data, deep learning technology can deeply understand and model the complexity of speech and its context relationship, with higher recognition accuracy and response speed. Compared with traditional systems, the voice assistant based on deep learning can recognize different accents and speaking speeds of users, and at the same time has stronger learning ability, and can improve its own performance and adaptability by continuously accumulating usage data. Summary of the Invention

[0006] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a method and device for a voice assistant system for online car-hailing based on deep learning to solve the problems raised in the above background art.

[0007] To achieve the above object, the present invention provides the following technical solution, a method for a voice assistant system for online car-hailing based on deep learning, specifically including the following steps:

[0008] Step S1: Randomly select online car-hailing order data from the online database, including the recordings of user voice commands and the corresponding text conversions, and perform feature engineering processing on the collected voice command data to generate a training set, a test set, and a validation set;

[0009] Step S2: Repeat Step S1 for data processing to generate multiple sets of order-dimensional data, load the training set and the test set, and perform multiple rounds of training until the loss function converges to generate an autoencoder model;

[0010] Step S3: Read the validation set data, input it into the autoencoder model, verify the performance of the model on the validation set, and analyze the output results;

[0011] Step S4: Repeat Step S2 and Step S3 to generate an optimal autoencoder model, input the real-time processed data into the trained autoencoder model, obtain the voice recognition and understanding results based on each order, and automatically process the orders according to the output results of the model.

[0012] In a preferred embodiment, in Step S1, randomly select online car-hailing order data from the online database, including the recordings of user voice commands and the corresponding text conversions, and perform feature engineering processing on the collected voice command data to generate a training set, a test set, and a validation set. The specific steps are as follows:

[0013] Step A1: Data collection: Identify and connect to the online database containing online car-hailing orders, use the method of random sampling to select a certain number of order data from the database to ensure sample representativeness, extract the recordings of user voice commands and their corresponding text conversions from the selected orders, and process the extracted voice commands and text data and store them as structured data. The voice command processing includes noise removal, audio format standardization, and sampling rate setting; the text data cleaning includes removing redundant spaces, punctuation marks, and normalizing vocabulary;

[0014] Step A2: Feature extraction: Extract audio features and text features. The text feature extraction is to convert the cleaned text commands into numerical features, and use the bag-of-words model to convert the text into a feature vector, representing the occurrence frequency of each word in the text;

[0015] Step A3: Feature combination: Combine the extracted audio features and text features to form a final feature vector, ensure that each order can generate a set of consistent feature vectors for subsequent model training, and divide the generated feature vector dataset into a training set, a test set, and a validation set.

[0016] In a preferred embodiment, in the step A2 of feature extraction, audio features and text features are extracted, and the audio feature extraction further includes the following steps:

[0017] Step A201, cutting and framing: The voice command x(t) is segmented into segments x n (t), and the length of each segment is 20 - 50 milliseconds to better capture features, and a window function is applied to each audio segment to reduce edge effects. The windowed segment is represented as: x' n (t) = x n (t)·ω(t); where x' n (t) is the windowed segment, x n (t) represents the segment after the voice command is segmented, is the window function, and the duration of each segment is T b = 20ms to 50ms;

[0018] Step A202, performing a short-time Fourier transform on each windowed segment to obtain a spectrum where X n (f,t) is the representation of the nth segment in the frequency domain, f is the frequency, t is the time, x' n (τ) is the time-domain signal of the nth segment after being processed by the window function, τ is the time variable, τ - t represents the deviation of the center position of the window function from the time t, and j is the imaginary unit;

[0019] Using a Mel filter bank to convolve the spectrum |X n (f,t)| 2 to obtain a Mel spectrum where H m (f) is the response of the mth Mel filter, F is the set of frequencies, and f is the frequency;

[0020] Performing a transform through a discrete cosine transform to obtain Mel frequency cepstral coefficients as where k is the index of the Mel frequency cepstrum, M is the number of Mel filters, and L n (m,t) is the logarithm of the Mel spectrum.

[0021] In a preferred embodiment, in the step S2, step S1 is repeated to perform data processing to generate multiple sets of order dimension data, and a training set and a test set are loaded for multiple rounds of training until the loss function converges to generate an autoencoder model. The specific steps are as follows:

[0022] Step B1. Autoencoder structure: Load the training set and the test set, generate a data iterator for step-by-step calling of the iterative model, initialize the model and related hyperparameters, including the input layer, hidden layer, output layer, encoding layer, decoding layer, gradient propagation, additional gradient, loss function, number of training rounds, and the size of each training data. Map the input α to the hidden layer representation h, where h = σ(W o α) + b o , where W o and b o are the weight and bias of the encoder respectively, and σ is the activation function; restore the hidden layer representation h back to the input space W d and b d are the weight and bias of the decoder respectively;

[0023] Step B2. Multi-round training: Use the reconstruction loss as the training objective, update the model parameters by backpropagation through the optimization algorithm until the loss function L reaches the maximum number of training rounds. After training, save the autoencoder model, check the reconstruction error to ensure that the model has good generalization ability. The specific calculation formula is as follows:

[0024]

[0025] where N is the number of data samples, α(i) is the true input, is the reconstructed output, and L is the loss function.

[0026] In a preferred embodiment, in step S3, read the validation set data and input it into the autoencoder model to verify the effect of the model on the validation set and analyze the output results. The specific steps are as follows:

[0027] Step C1. Read the validation set data from storage, use the autoencoder model to infer the validation set data, generate the reconstructed output, and store the reconstructed output in a new variable. Evaluate the reconstruction ability of the autoencoder model through the mean square error;

[0028] Step C2. Analyze the output results: Compare the original validation data with the reconstructed data, visualize it by plotting a waveform diagram, and plot the original signal and the reconstructed signal in the same graph to observe their similarity.

[0029] In a preferred embodiment, in step S4, repeat steps S2 and S3 to generate the optimal autoencoder model, input the real-time processed data into the trained autoencoder model, obtain the speech recognition and understanding results based on each order, and automatically process the order according to the output results of the model. The specific steps are as follows:

[0030] Step D1: Set up an interface to receive real-time order voice data, ensure that the data can be transmitted in a timely manner and is ready for processing. Process the real-time received voice data to ensure that its format is consistent with the data format of the trained model. Input the real-time processed data into the trained autoencoder model and obtain its output;

[0031] Step D2: Analyze the output of the autoencoder, extract relevant voice instructions, and automatically perform relevant operations according to the extracted order information, including creating new orders, modifying orders, and sending notifications.

[0032] This application also provides a device for a network car-hailing voice assistant system based on deep learning, which specifically includes a data processing module, a model training module, a model verification module, and an order management module;

[0033] Data processing module: Randomly select network car-hailing order data from an online database, including recordings of user voice instructions and corresponding text conversions, and perform feature engineering processing on the collected voice instruction data to generate a training set, a test set, and a validation set;

[0034] Model training module: Load the training set and the test set, conduct multiple rounds of training until the loss function converges, and generate an autoencoder model;

[0035] Model verification module: Read the validation set data and input it into the autoencoder model to verify the effect of the model on the validation set and analyze the output results;

[0036] Order management module: Input the real-time processed data into the trained autoencoder model to obtain the voice recognition and understanding results based on each order, and automatically process the order according to the output results of the model.

[0037] The beneficial effects of the present invention are as follows: Randomly select network car-hailing order data from an online database, including recordings of user voice instructions and corresponding text conversions, and perform feature engineering processing on the collected voice instruction data to generate a training set, a test set, and a validation set. Load the training set and the test set, conduct multiple rounds of training until the loss function converges, and generate an autoencoder model. Read the validation set data and input it into the autoencoder model to verify the effect of the model on the validation set and analyze the output results. Input the real-time processed data into the trained autoencoder model to obtain the voice recognition and understanding results based on each order, and automatically process the order according to the output results of the model. In network car-hailing services, the present invention can more accurately and realistically understand users' voice instructions and provide more convenient, efficient, and user-friendly services, which are incomparable to traditional voice recognition and natural language processing technologies. Description of the Drawings

[0038] Figure 1This is the method flow chart of the present invention. Detailed implementation manners

[0039] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0040] In the description of the present application, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present application, "a plurality" means two or more, unless otherwise specifically defined.

[0041] In the description of the present application, the term "for example" is used to mean "serving as an example, illustration, or explanation". Any embodiment described as "for example" in the present application is not necessarily construed as being more preferred or more advantageous than other embodiments. In order for any person skilled in the art to implement and use the present invention, the following description is given. In the following description, details are set forth for the purpose of explanation. It should be understood that those of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other instances, well-known structures and processes are not elaborated in detail to avoid unnecessary details from obscuring the description of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope that conforms to the principles and features disclosed in the present application.

[0042] Embodiment 1

[0043] This embodiment provides a method for a online car-hailing voice assistant system based on deep learning as shown in Figure 1 and specifically includes the following steps:

[0044] Step S1: Randomly select online car-hailing order data from an online database, including recordings of user voice instructions and corresponding text conversions, and perform feature engineering processing on the collected voice instruction data to generate a training set, a test set, and a validation set;

[0045] Step S2: Repeat Step S1 to perform data processing to generate multiple groups of order dimension data, load the training set and the test set, and perform multiple rounds of training until the loss function converges to generate an autoencoder model;

[0046] Step S3: Read the validation set data, input it into the autoencoder model, verify the model's performance on the validation set, and analyze the output results;

[0047] Step S4: Repeat Step S2 and Step S3 to generate the optimal autoencoder model, input the real-time processed data into the trained autoencoder model, obtain the speech recognition and understanding results based on each order, and automatically process the orders according to the model's output results.

[0048] Preferably, in Step S1, randomly select online car-hailing order data from the online database, including the recordings of user voice commands and the corresponding text conversions, and perform feature engineering processing on the collected voice command data to generate a training set, a test set, and a validation set, improve the model's performance on unseen data, and enhance the model's generalization ability. The specific steps are as follows:

[0049] Step A1: Data collection: Identify and connect to the online database containing online car-hailing orders, use the method of random sampling to select a certain number of order data from the database to ensure sample representativeness, extract the recordings of user voice commands and their corresponding text conversions from the selected orders, and process the extracted voice commands and text data and store them as structured data. The voice command processing includes removing noise, standardizing the audio format, and setting the sampling rate; the text data cleaning includes removing redundant spaces, punctuation marks, and normalizing vocabulary;

[0050] Step A2: Feature extraction: Extract audio features and text features. The text feature extraction is to convert the cleaned text commands into numerical features, use the bag-of-words model to convert the text into feature vectors, representing the occurrence frequency of each word in the text. By extracting features from the voice data, the data dimension can be reduced, the complexity of model training can be reduced, the training speed can be accelerated, and the model can converge more easily;

[0051] Step A3: Merge features: Merge the extracted audio features and text features to form the final feature vector, ensure that each order can generate a set of consistent feature vectors for subsequent model training, and divide the generated feature vector dataset into a training set, a test set, and a validation set.

[0052] Preferably, in the feature extraction of Step A2, extract audio features and text features, and the audio feature extraction further includes the following steps:

[0053] Step A201: Cutting and framing: Divide the voice command x(t) into segments x n (t), and the length of each segment is 20 - 50 milliseconds to better capture features, and apply a window function to each audio segment to reduce edge effects. The windowed segment is represented as: x'n x'(t) = x(t)·ω(t); where x' n (t) is the windowed segment, and x n (t) represents the segment after speech command segmentation, n ω(t) is the window function, and the duration of each segment is T = 20 ms to 50 ms; b

[0054] Step A202: Perform short-time Fourier transform on each windowed segment to obtain the spectrum where X n (f,t) is the representation of the nth segment in the frequency domain, f is the frequency, t is the time, and x' n (τ) is the time-domain signal of the nth segment after window function processing, τ is the time variable, τ - t represents the deviation of the center position of the window function from time t, and j is the imaginary unit;

[0055] Use the Mel filter bank to convolve the spectrum |X n (f,t)| 2 to obtain the Mel spectrum where H m (f) is the response of the mth Mel filter, F is the set of frequencies, and f is the frequency;

[0056] Perform transformation through discrete cosine transform to obtain the Mel-frequency cepstral coefficients as where k is the index of the Mel-frequency cepstrum, M is the number of Mel filters, and L n (m,t) is the logarithm of the Mel spectrum.

[0057] Preferably, in step S2, repeat step S1 for data processing to generate multiple sets of order dimension data, load the training set and test set, and perform multiple rounds of training until the loss function converges to generate an autoencoder model. The specific steps are as follows:

[0058] Step B1: Autoencoder structure: Load the training set and test set, generate a data iterator for step-by-step calling of the iterative model, initialize the model and related hyperparameters, including the input layer, hidden layer, output layer, encoding layer, decoding layer, gradient propagation, additional gradient, loss function, number of training rounds, and size of each training data. Map the input α to the hidden layer representation h as h = σ(W o α) + b o where W o and b o are the weights and biases of the encoder respectively, and σ is the activation function; restore the hidden layer representation h back to the input space W d and b d ​They are the weights and biases of the decoder respectively;

[0059] Step B2, multi-round training: Use the reconstruction loss as the training objective, update the model parameters through backpropagation of the optimization algorithm until the loss function L reaches the maximum number of training rounds. After the training is completed, save the autoencoder model, check the reconstruction error, and ensure that the model has good generalization ability. The specific calculation formula is as follows:

[0060]

[0061] Among them, N is the number of data samples, α(i) is the true input, is the reconstructed output, and L is the loss function.

[0062] Preferably, in step S3, read the validation set data, input it into the autoencoder model, verify the effect of the model on the validation set, and analyze the output results, which can evaluate the performance of the model, understand its performance on unseen data, and verify the effectiveness and accuracy of the model. The specific steps are as follows:

[0063] Step C1, read the validation set data from storage, use the autoencoder model to perform inference on the validation set data, generate the reconstructed output, and store the reconstructed output in a new variable. Evaluate the reconstruction ability of the autoencoder model through the mean square error;

[0064] Step C2, analyze the output results: Compare the original validation data with the reconstructed data, visualize it by plotting a waveform diagram, and plot the original signal and the reconstructed signal in the same graph to observe their similarity.

[0065] Preferably, in step S4, repeat steps S2 and S3 to generate the optimal autoencoder model, input the real-time processed data into the trained autoencoder model, obtain the speech recognition and understanding results based on each order, automatically process the order according to the output results of the model, reduce manual intervention, improve efficiency, reduce the error rate, and can greatly speed up the response time, making the user experience smoother. The specific steps are as follows:

[0066] Step D1, set up an interface to receive real-time order voice data, ensure that the data can be transmitted in a timely manner and is ready for processing, process the real-time received voice data to ensure that its format is consistent with the data format of the trained model, and input the real-time processed data into the trained autoencoder model to obtain its output;

[0067] Step D2: Analyze the output of the autoencoder, extract relevant voice commands, and automatically execute relevant operations according to the extracted order information, including creating new orders, modifying orders, and sending notifications. Through the output of the high-precision autoencoder, it is possible to accurately identify the user's voice commands, provide reliable data basis for order processing, and reduce misunderstandings and errors.

[0068] Embodiment 2

[0069] This embodiment provides a device for an online car-hailing voice assistant system based on deep learning, specifically including a data processing module, a model training module, a model verification module, and an order management module;

[0070] Data processing module: Randomly select online car-hailing order data from the online database, including recordings of user voice commands and corresponding text conversions, and perform feature engineering processing on the collected voice command data to generate a training set, a test set, and a validation set;

[0071] Model training module: Load the training set and the test set, perform multiple rounds of training until the loss function converges, and generate an autoencoder model;

[0072] Model verification module: Read the validation set data and input it into the autoencoder model to verify the effect of the model on the validation set and analyze the output results;

[0073] Order management module: Input the real-time processed data into the trained autoencoder model to obtain the voice recognition and understanding results based on each order, and automatically process the orders according to the output results of the model.

[0074] It should be noted that in the above embodiments, the descriptions of each embodiment have their own focuses. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0075] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0076] The present invention will be described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or a means for implementing the functions specified in multiple blocks.

[0077] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means, and the instruction means implements the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or a means for implementing the functions specified in multiple blocks.

[0078] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or a means for implementing the functions specified in multiple blocks.

[0079] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic inventive concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present invention.

[0080] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A method for an online car-hailing voice assistant system based on deep learning, characterized in that, Specifically, it includes the following steps: Step S1: Randomly select online car-hailing order data from the online database, including the recordings of user voice commands and the corresponding text conversions, and perform feature engineering processing on the collected voice command data to generate a training set, a test set, and a validation set; Step S2: Repeat Step S1 to perform data processing to generate multiple sets of order-dimensional data, load the training set and the test set, and perform multiple rounds of training until the loss function converges to generate an autoencoder model; Step S3: Read the validation set data and input it into the autoencoder model to verify the effect of the model on the validation set and analyze the output results; Step S4: Repeat Step S2 and Step S3 to generate an optimal autoencoder model, input the real-time processed data into the trained autoencoder model, obtain the speech recognition and understanding results based on each order, and automatically process the order according to the output results of the model.

2. The method of an online car-hailing voice assistant system based on deep learning according to claim 1, characterized in that: In Step S1, randomly select online car-hailing order data from the online database, including the recordings of user voice commands and the corresponding text conversions, and perform feature engineering processing on the collected voice command data to generate a training set, a test set, and a validation set. The specific steps are as follows: Step A1: Data collection: Identify and connect to the online database containing online car-hailing orders, use the method of random sampling to select order data from the database, extract the recordings of user voice commands and their corresponding text conversions from the selected orders, and process the extracted voice commands and text data and store them as structured data; Step A2: Feature extraction: Extract audio features and text features. The text feature extraction is to convert the cleaned text commands into numerical features, and use the bag-of-words model to convert the text into a feature vector, indicating the frequency of each word in the text; Step A3: Merge features: Merge the extracted audio features and text features to form a final feature vector, and divide the generated feature vector dataset into a training set, a test set, and a validation set.

3. The method of an online car-hailing voice assistant system based on deep learning according to claim 2, wherein: In the feature extraction of Step A2, extract audio features and text features. The audio feature extraction further includes the following steps: Step A201, Cutting and Framing: Split the voice command x(t) into segments x n (t), where the length of each segment is 20 - 50 milliseconds to better capture features, and apply a window function to each audio segment to reduce edge effects. The windowed segment is represented as: x' n (t) = x n (t)·ω(t); where x' n (t) is the windowed segment, x n (t) represents the segment after splitting the voice command, ω(t) is the window function, and the duration of each segment is T b = 20ms to 50ms; Step A202: Perform short-time Fourier transform on each windowed segment to obtain the spectrum where X n (f, t) is the representation of the nth segment in the frequency domain, f is the frequency, t is the time, and x' n (τ) is the time-domain signal of the nth segment after being processed by the window function, τ is the time variable, τ - t represents the deviation between the center position of the window function and the time t, and j is the imaginary unit; Convolve the spectrum |X n (f,t)| 2 with the Mel filter bank to obtain the Mel spectrum where H m (f) is the response of the m-th Mel filter, F is the set of frequencies, and f is the frequency; Transform through discrete cosine transform to obtain Mel frequency cepstral coefficients as where k is the index of the Mel frequency cepstrum, M is the number of Mel filters, and L n (m,t) is the logarithm of the Mel spectrum.

4. The method of an online car-hailing voice assistant system based on deep learning according to claim 1, wherein: In Step S2, repeat Step S1 to perform data processing to generate multiple sets of order-dimensional data, load the training set and the test set, and perform multiple rounds of training until the loss function converges to generate an autoencoder model. The specific steps are as follows: Step B1. Autoencoder structure: Load the training set and the test set, generate a data iterator for step-by-step calling of the iterative model, initialize the model and related hyperparameters, and map the input α to the hidden-layer representation h as h = σ(W o α) + b o , where W o and b o are the weights and biases of the encoder respectively, and σ is the activation function; restore the hidden-layer representation h back to the input space W d and b d are the weights and biases of the decoder respectively; Step B2: Multiple rounds of training: Use the reconstruction loss as the training objective, update the model parameters by backpropagation through the optimization algorithm until the loss function L reaches the maximum number of training rounds, and save the autoencoder model after the training ends.

5. The method of an online car-hailing voice assistant system based on deep learning according to claim 4, characterized in that: The specific calculation formula of the loss function L is as follows: where N is the number of data samples, α(i) is the true input, is the reconstructed output, and L is the loss function.

6. The method of an online car-hailing voice assistant system based on deep learning according to claim 1, characterized in that: In Step S3, read the validation set data and input it into the autoencoder model to verify the effect of the model on the validation set and analyze the output results. The specific steps are as follows: Step C1: Read the validation set data from the storage, use the autoencoder model to perform inference on the validation set data to generate a reconstructed output, store the reconstructed output in a new variable, and evaluate the reconstruction ability of the autoencoder model through the mean squared error; Step C2, Analyze the output results: Compare the original verification data with the reconstructed data, visualize it by plotting a waveform diagram, and plot the original signal and the reconstructed signal in the same graph to observe the similarity between the two.

7. The method of a ride-hailing voice assistant system based on deep learning according to claim 1, characterized in that: In the said step S4, repeat step S2 and step S3 to generate an optimal autoencoder model, input the real-time processed data into the trained autoencoder model, obtain the speech recognition and understanding results based on each order, and automatically process the order according to the output results of the model. The specific steps are as follows: Step D1, Set up an interface to receive real-time order voice data, process the real-time received voice data, input the real-time processed data into the trained autoencoder model, and obtain its output; Step D2, Analyze the output of the autoencoder, extract relevant voice instructions, and automatically execute relevant operations according to the extracted order information, including creating a new order, modifying an order, and sending a notification.

8. An apparatus of a network car-hailing voice assistant system based on deep learning is applied to the method of a network car-hailing voice assistant system based on deep learning according to any one of claims 1-7, characterized in that: It includes a data processing module, a model training module, a model verification module, and an order management module; Data processing module: Randomly select online car-hailing order data from the online database, including recordings of user voice instructions and corresponding text conversions, and perform feature engineering processing on the collected voice instruction data to generate a training set, a test set, and a validation set; Model training module: Load the training set and the test set, perform multiple rounds of training until the loss function converges, and generate an autoencoder model; Model verification module: Read the validation set data, input it into the autoencoder model, verify the effect of the model on the validation set, and analyze the output results; Order management module: Input the real-time processed data into the trained autoencoder model, obtain the speech recognition and understanding results based on each order, and automatically process the order according to the output results of the model.