Driving control system and method based on voice recognition

Through a speech recognition system based on deep learning and reinforcement learning, the complex operation and inefficiency of the driving crane control system are solved, precise and rapid driving control is achieved, and the accuracy and stability of driving control is improved.

CN120279911AActive Publication Date: 2025-07-08CHENGDU EDERLES TECH CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510758107.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-08
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

The existing driving crane control system relies on manual operation or simple remote control operation, and has problems such as complex operation and low efficiency. The voice recognition technology is insufficient in recognition accuracy, anti-interference ability and system stability.

Method used

A voice recognition system based on deep learning and reinforcement learning is adopted. Voice command data is collected through a microphone array, and after denoising, a deep learning algorithm is used to identify text command data, and driving control parameters are determined through reinforcement learning algorithms, ultimately realizing accurate and fast driving control.

Benefits of technology

It realizes accurate and fast driving control, improves operating efficiency, is suitable for a variety of scenarios, and improves the accuracy and stability of driving control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279911A_ABST
    Figure CN120279911A_ABST
Patent Text Reader

Abstract

The invention discloses a driving control system and method based on voice recognition, and belongs to the technical field of control. Voice instruction data of an operator are collected through a microphone array, and after the voice instruction data are denoised, the denoised voice instruction data are obtained; then, a deep learning algorithm is adopted to recognize the de-noised voice instruction data, determine character instruction data, analyze the character instruction data, determine a driving control keyword, convert the driving control keyword into a word vector, and finally, on the basis of the word vector and the driving state information, the driving state information of the vehicle is obtained. The crane control parameters are determined by adopting the reinforcement learning algorithm, the crane is controlled to work on the basis of the crane control parameters, accurate and rapid control over the crane can be achieved, the method is suitable for many scenes, and the crane control efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of control technology, and particularly relates to a vehicle control system and method based on speech recognition. Background Art

[0002] A vehicle generally refers to an overhead crane, also known as a bridge crane, which is horizontally installed above workshops, warehouses, and open storage yards for lifting materials at a certain height. It consists of a trolley (bridge) and a crab. The trolley runs longitudinally along the tracks laid on both sides of the elevated structure, and the crab runs transversely on the tracks of the trolley, forming a rectangular working range. The overhead crane has a large lifting capacity and a high working level, and is widely used in fields such as mechanical manufacturing, metallurgy, and logistics. It can be equipped with different lifting devices, such as hooks, grabs, electromagnetic chucks, etc., to meet different material handling requirements. Compared with other types of cranes, the overhead crane has the advantages of a wide coverage area, high efficiency, and flexible operation, and is an important and indispensable device in modern production. As an important material handling device in industrial production, the safety and efficiency of its operation are crucial. The traditional control system of overhead cranes mainly relies on manual operation or simple remote control operation, which has problems such as complex operation, low efficiency, and potential safety hazards. In recent years, with the rapid development of speech recognition technology, its application in the field of industrial control has gradually attracted attention. However, the existing vehicle control systems based on speech recognition for overhead cranes still have deficiencies in terms of recognition accuracy, anti-interference ability, and system stability. Summary of the Invention

[0003] The present invention provides a vehicle control system and method based on speech recognition to solve the problems of complex operation and low efficiency in the traditional control system of overhead cranes in the prior art, which mainly relies on manual operation or simple remote control operation.

[0004] On the one hand, the present invention provides a vehicle control system based on speech recognition, including: A speech acquisition module, configured to collect the voice command data of an operator through a microphone array, and obtain the denoised voice command data after denoising the voice command data; A speech recognition module, configured to recognize the denoised voice command data by using a deep learning algorithm to determine the text command data; An instruction parsing module, configured to parse the text command data to determine vehicle control keywords, and convert the vehicle control keywords into word vectors; A reinforcement learning module, configured to determine vehicle control parameters by using a reinforcement learning algorithm based on the word vectors and the state information of the vehicle; A driving control module, configured to control a vehicle to perform operations based on the driving control parameters, and complete driving control based on voice recognition.

[0005] In some possible implementation manners, the voice recognition module includes a voice recognition model deployment sub-module, a voice data feature extraction sub-module, and a voice feature recognition sub-module; The voice recognition model deployment sub-module is configured to pre-deploy a voice recognition model for voice recognition by using a deep learning algorithm; The voice data feature extraction sub-module is configured to extract features from the denoised voice command data to obtain voice features; The voice feature recognition sub-module is configured to schedule the voice recognition model pre-deployed by the voice recognition model deployment sub-module to recognize the voice features and determine text command data.

[0006] In some possible implementation manners, pre-deploying a voice recognition model for voice recognition by using a deep learning algorithm includes: Initializing hyperparameters of the deep learning algorithm by using a chaotic mapping sequence to determine a plurality of different parameter vectors; wherein, the parameter vector includes all or part of the hyperparameters to be optimized of the deep learning algorithm; Obtaining the loss function value of each parameter vector, and determining the fitness corresponding to the parameter vector according to the loss function value; Determining an optimal parameter vector and a worst parameter vector according to the fitness corresponding to the parameter vector; Based on the optimal parameter vector, performing an initial search on the parameter vector by using a variable spiral position selection mechanism to determine the parameter vector after the initial search; For the parameter vector after the initial search, performing a local and global balance search on the parameter vector by using an adaptive probability response mechanism according to the worst parameter vector to determine the parameter vector after the balance search; Judging whether the current training times is greater than or equal to a preset maximum training times. If so, re-determining the optimal parameter vector according to the parameter vector after the balance search, otherwise returning to the step of obtaining the fitness; Determining the final hyperparameters of the deep learning algorithm according to the re-determined optimal parameter vector, and deploying a voice recognition model for voice recognition according to the final hyperparameters of the deep learning algorithm.

[0007] In some possible implementation manners, based on the optimal parameter vector, performing an initial search on the parameter vector by using a variable spiral position selection mechanism to determine the parameter vector after the initial search, includes: Based on the current training times, determining an adaptive inertia weight by using a hyperbolic tangent function; The optimal parameter vector is weighted using the adaptive inertia weight to determine the weighted optimal parameter vector; For any parameter vector, a first difference vector between the optimal parameter vector and the parameter vector is obtained; For any parameter vector, the position of the parameter vector is transformed using a variable spiral function to determine the parameter vector after the position transformation; Based on the weighted optimal parameter vector, the first difference vector, and the parameter vector after the position transformation, the parameter vector after the initial search is determined.

[0008] In some possible implementation manners, for the parameter vector after the initial search, according to the worst parameter vector, an adaptive probability response mechanism is used to perform a local and global balanced search on the parameter vector to determine the parameter vector after the balanced search, including: The fitness corresponding to each parameter vector after the initial search is obtained, and the population dispersion degree is obtained according to the fitness corresponding to each parameter vector after the initial search; Based on the population dispersion degree, a local development influence factor and a global development influence factor are obtained; The single search rate is obtained according to the fitness corresponding to each parameter vector after the initial search, and the average search rate during multiple training processes is determined according to the single search rate; Based on the average search speed, a local development threshold parameter and a global development threshold parameter are determined; Based on the local development influence factor and the local development threshold parameter, the local development probability is determined, and based on the local development probability, a position transfer strategy is used for local development to determine the parameter vector after local development; Based on the global development influence factor and the global development threshold parameter, the global development probability is determined, and based on the global development probability, a chaotic global search strategy is used for global development to determine the parameter vector after global development; For any parameter vector after the initial search, if the parameter vector has not undergone global development and local development, the original parameter vector after the initial search is directly used as the parameter vector after the balanced search; If the parameter vector has undergone global development or local development, the parameter vector after global development or the parameter vector after local development is used as the parameter vector after the balanced search; If the parameter vector has undergone both global development and local development, an annealing simulation algorithm is used to select the parameter vector after global development or the parameter vector after local development as the parameter vector after the balanced search.

[0009] In some possible embodiments, according to the local development probability, a position transfer strategy is adopted for local development to determine the parameter vector after local development, including: Randomly generate a first decision factor within the interval (0, 1), and determine whether the first decision factor is less than the local development probability. If so, adopt the position transfer strategy for local development to determine the parameter vector after local development; otherwise, do not perform local development. Adopting the position transfer strategy for local development to determine the parameter vector after local development includes: For any parameter vector after initial search, determine the transfer probabilities corresponding to other parameter vectors according to the fitness corresponding to the parameter vector after initial search. Determine the target cooperation vector for the parameter vector after initial search according to the transfer probabilities corresponding to other parameter vectors. Perform local development on the parameter vector after initial search according to the target cooperation vector to determine the parameter vector after local development.

[0010] In some possible embodiments, according to the global development probability, a chaotic global search strategy is adopted for global development to determine the parameter vector after global development, including: Randomly generate a second decision factor within the interval (0, 1), and determine whether the second decision factor is less than the global development probability. If so, adopt the chaotic global search strategy for global development to determine the parameter vector after global development; otherwise, do not perform global development. Adopting the chaotic global search strategy for global development to determine the parameter vector after global development includes: Obtain the chaotic mapping factor in the current training process. Determine the search center vector according to the search upper limit vector and the search lower limit vector. Perform global development on the parameter vector after initial search according to the chaotic mapping factor in combination with the sine function to determine the parameter vector after global development.

[0011] In some possible embodiments, perform feature extraction on the denoised voice command data to obtain voice features, including: extracting the MFCC of the denoised voice command data to obtain voice features.

[0012] In some possible embodiments, based on the word vector and the driving state information, adopt a reinforcement learning algorithm to determine the driving control parameters, including: Collect the driving state information; wherein, the state information includes position, attitude, and speed. Construct the word vector and the driving state information into the current state space, and use the current state space as the input of the reinforcement learning algorithm to determine the selected action and obtain the driving control parameters; Among them, the action is selected from the action space, each action represents a driving control parameter, and the action space should include all driving control parameters.

[0013] On the other hand, the present invention provides a driving control method based on voice recognition, including: Collect the voice command data of the operator through the microphone array, and after denoising the voice command data, obtain the denoised voice command data; Use a deep learning algorithm to recognize the denoised voice command data to determine the text command data; Parse the text command data to determine the driving control keywords, and convert the driving control keywords into word vectors; Based on the word vector and the driving state information, use a reinforcement learning algorithm to determine the driving control parameters; Based on the driving control parameters, control the driving to perform operations to complete the driving control based on voice recognition.

[0014] A driving control system and method based on voice recognition provided by the present invention collect the voice command data of the operator through the microphone array, and after denoising the voice command data, obtain the denoised voice command data, then use a deep learning algorithm to recognize the denoised voice command data to determine the text command data, and parse the text command data to determine the driving control keywords, and convert the driving control keywords into word vectors. Finally, based on the word vector and the driving state information, use a reinforcement learning algorithm to determine the driving control parameters, and based on the driving control parameters, control the driving to perform operations, which can achieve precise and rapid control of the driving, is applicable to many scenarios, and improves the driving control efficiency. Description of the Drawings

[0015] The drawings here are incorporated into the specification and form a part of the specification, showing the embodiments in line with the present invention, and are used together with the specification to explain the principles of the present invention.

[0016] Figure 1 It is a schematic structural diagram of a driving control system based on voice recognition provided by an embodiment of the present invention.

[0017] Figure 2 It is a schematic flow diagram of a driving control method based on voice recognition provided by an embodiment of the present invention.

[0018] Among them, 101 - voice acquisition module, 102 - voice recognition module, 103 - instruction parsing module, 104 - reinforcement learning module, 105 - driving control module.

[0019] Through the above - mentioned drawings, specific embodiments of the present invention have been shown, and there will be a more detailed description hereinafter. These drawings and written descriptions are not intended to limit the scope of the inventive concept in any way, but to illustrate the concept of the present invention to those skilled in the art by referring to specific embodiments. Detailed implementation manners

[0020] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0021] The embodiments of the present invention will be described in detail below with reference to the drawings.

[0022] As Figure 1 shown, an embodiment of the present invention provides a driving control system based on voice recognition, including: A voice acquisition module 101, configured to collect voice command data of an operator through a microphone array, and after denoising the voice command data, obtain the denoised voice command data; A voice recognition module 102, configured to use a deep learning algorithm to recognize the denoised voice command data to determine text command data; The voice features of the denoised voice command data can be extracted first, and then the voice features can be recognized to determine the text command data.

[0023] An instruction parsing module 103, configured to parse the text command data to determine driving control keywords, and convert the driving control keywords into word vectors; The driving control keywords refer to pre - set control keywords, such as forward, backward, down, up, and various distance digital data.

[0024] A reinforcement learning module 104, configured to determine driving control parameters based on the word vectors and the state information of the vehicle using a reinforcement learning algorithm; In the case of aging of the traveling motor or inaccurate command from the commander (such as not stating the forward distance), the traveling crane may not be able to reach the operation position accurately. However, in the embodiments of the present invention, the control parameters of the traveling crane are determined through a reinforcement learning algorithm, which can not only achieve effective control of the traveling crane, but also learn long-term habits to select more accurate control parameters of the traveling crane, enabling it to accurately reach the operation position in a fixed operation scenario and effectively improving the operation efficiency of the traveling crane.

[0025] The traveling crane control module 105 is configured to control the traveling crane to perform operations based on the traveling crane control parameters, thereby completing the traveling crane control based on speech recognition.

[0026] In some possible implementation manners, the speech recognition module includes a speech recognition model deployment sub-module, a speech data feature extraction sub-module, and a speech feature recognition sub-module; The speech recognition model deployment sub-module is configured to pre-deploy a speech recognition model for speech recognition by using a deep learning algorithm; The speech data feature extraction sub-module is configured to extract features from the denoised speech command data to obtain speech features; denoising is a relatively conventional technical means and will not be elaborated in the embodiments of the present invention.

[0027] The speech feature recognition sub-module is configured to schedule the speech recognition model pre-deployed by the speech recognition model deployment sub-module to recognize the speech features and determine the text command data.

[0028] Generally, the deep learning algorithm cannot be directly used for speech recognition. It is necessary to first learn the historical data to obtain a speech recognition model for speech recognition. In the traditional technology, the gradient descent method is generally used to update the hyperparameters of the deep learning algorithm, which will cause the search process to fall into a local optimum, and finally the trained speech recognition model cannot effectively perform speech recognition. Therefore, the embodiments of the present invention provide the following pre-deployment algorithm to achieve global training while ensuring the training speed and training accuracy, and finally ensuring the accuracy of speech recognition.

[0029] In some possible implementation manners, pre-deploying a speech recognition model for speech recognition by using a deep learning algorithm includes: A1. Initializing the hyperparameters of the deep learning algorithm by using a chaotic mapping sequence to determine a plurality of different parameter vectors; wherein, the parameter vector includes all or part of the hyperparameters to be optimized of the deep learning algorithm, The above step A1 indicates that all or part of the hyperparameters of the deep learning algorithm can be optimized during the training process, so that different optimization objectives can be set according to actual requirements.

[0030] The deep learning algorithm can be set as: the CNN-CTC algorithm; among them, CTC (Connectionist Temporal Classification) adopts a learning algorithm that directly calculates the posterior probability of the overall output of the sequence, and can complete the conversion of a speech signal into a corresponding word sequence or other text unit sequences with only one neural network. Under the CTC end-to-end framework, the convolutional neural network has greater application potential. CTC can splice all the feature vectors of the entire speech signal into a large feature map as the input. Through the stacking of the convolutional layer and pooling layer of the CNN (Convolutional Neural Network), the "receptive field" of the neurons in the output layer is improved as the number of layers deepens, so that the final output of the network can capture more context information between speech frames. Secondly, in the selection of the modeling unit, since CTC introduces the "blank" label and does not need to strictly divide the category of each speech frame, larger-grained modeling units such as syllables and Chinese characters can be selected. Moreover, the unique convolution and downsampling operations of the CNN not only compress the input feature map but also integrate and screen the speech frame data, and can extract more essential features from the speech signal with larger-grained modeling units containing more complex information.

[0031] Initializing the hyperparameters of the deep learning algorithm using a chaotic mapping sequence may include: For each hyperparameter of the deep learning algorithm, it can be randomly initialized between the upper and lower limits of the hyperparameter (for example, the upper limit of the hyperparameter in the embodiment of the present invention is generally 1, and the lower limit is generally 0), and the hyperparameter after initialization is encoded as a vector to obtain a basic parameter vector; Based on the basic parameter vector, a chaotic mapping sequence is used to obtain multiple different parameter vectors as: Among them, represents the i th d -dimensional hyperparameter of the d th parameter vector in the initialization process, i = 1, 2,..., D, D represents the total dimension of the hyperparameters, and when i = 1, i represents the d -dimensional hyperparameter of the th basic parameter vector in the initialization process, i represents the d -dimensional hyperparameter of the rd parameter vector in the initialization process, represents the second chaotic mapping parameter and is set to 0.3; represents the sine function, represents the pi, represents the first random number between the interval (0, 1), represents the second random number between the interval (0, 1), represents the third random number between the interval (0, 1), represents the fourth random number between the interval (0, 1).

[0032] Compared with the existing Tent chaotic mapping, the chaotic mapping method provided by the embodiments of the present invention can make the parameter vectors in the population be extremely evenly distributed in the solution space, and the uniform distribution helps to ensure that the relative distances between the parameter vectors are large, thereby prompting the algorithm to better explore the search space. This helps to accelerate the convergence process of the algorithm and improve the possibility of finding a solution within a limited number of iterations.

[0033] A2. Obtain the loss function value of each parameter vector, and determine the fitness corresponding to the parameter vector according to the loss function value; For example, the loss function value can be added to a preset constant term (such as 0.001), and then the reciprocal is taken to obtain the fitness of the parameter vector.

[0034] A3. Determine the optimal parameter vector and the worst parameter vector according to the fitness corresponding to the parameter vector; The larger the fitness, the closer the parameter vector is to the optimal solution in the solution space. Therefore, the parameter vector with the largest fitness is the optimal parameter vector, and the parameter vector with the smallest fitness is the worst parameter vector.

[0035] A4. Based on the optimal parameter vector, use the variable spiral position selection mechanism to initialize the search for the parameter vector and determine the parameter vector after the initialization search; A5. For the parameter vector after the initialization search, according to the worst parameter vector, use the adaptive probability response mechanism to perform local and global balanced search on the parameter vector and determine the parameter vector after the balanced search; A6. Judge whether the current training times are greater than or equal to the preset maximum training times. If so, re-determine the optimal parameter vector according to the parameter vector after the balanced search, otherwise return to the step of obtaining the fitness; A7. Determine the final hyperparameters of the deep learning algorithm according to the re-determined optimal parameter vector, and deploy a speech recognition model for speech recognition according to the final hyperparameters of the deep learning algorithm.

[0036] Optionally, after each search of the parameter vector, boundary crossing processing can be performed on the parameter vector to ensure the validity of the parameter vector.

[0037] In some possible implementation manners, based on the optimal parameter vector, a variable spiral position selection mechanism is adopted to perform an initial search on the parameter vector to determine the parameter vector after the initial search, including: Based on the current number of training times, the hyperbolic tangent function is used to determine the adaptive inertia weight as: Wherein, represents the adaptive inertia weight, represents the maximum value of the adaptive inertia weight, represents the minimum value of the adaptive inertia weight, represents the hyperbolic tangent function, represents the maximum number of training times, represents the current number of training times; In order to reduce the probability of the algorithm entering the local optimal solution, the hyperbolic tangent function is introduced into the inertia weight provided by the embodiments of the present invention, which not only avoids the algorithm from entering the local optimal state too early in the early stage of iteration, but also strengthens the local search ability of the algorithm in the middle and late stages, can more accurately find the global optimal solution, and balance the global and local search abilities of the population.

[0038] The optimal parameter vector is weighted by using the adaptive inertia weight to determine the weighted optimal parameter vector as: ; wherein, represents the optimal parameter vector in the t th training process; For any parameter vector, the first difference vector between the optimal parameter vector and the parameter vector is obtained as: ; wherein, represents the t th parameter vector in the j th training process, j = 1, 2,..., NP, NP represents the total number of parameter vectors, and NP is an even number; For any parameter vector, the variable spiral function is used to perform a position transformation on the parameter vector to determine the parameter vector after the position transformation as: ; wherein, represents the natural constant, cos represents the cosine function, q represents the spiral shape constant, l represents a random variable spiral control factor between (-1, 1); Based on the weighted optimal parameter vector, the first difference vector, and the parameter vector after position transformation, determine the parameter vector after initial search as: ; where represents the j th parameter vector after initial search, represents the fifth random number between the interval (0, 1); The embodiment of the present invention adopts a variable spiral position selection mechanism to perform an initial search on the parameter vector. Based on the position of the optimal parameter vector, combined with the inertia weight and variable spiral search, it can have a larger search range in the initial stage of the algorithm, and gradually search for a finer solution around the optimal position in the later stage of the search, effectively improving the training performance of the algorithm.

[0039] In some possible implementation manners, for the parameter vector after initial search, according to the worst parameter vector, an adaptive probability response mechanism is adopted to perform a local and global balanced search on the parameter vector to determine the parameter vector after balanced search, including: Obtain the fitness corresponding to each parameter vector after initial search, and based on the fitness corresponding to each parameter vector after initial search, obtain the population dispersion degree as: ; where represents the population dispersion degree, represents the n th fitness corresponding to the parameter vector after initial search, represents the average fitness of the parameter vectors after initial search, n = 1, 2, …, NP; Based on the population dispersion degree, obtain the local development influence factor and the global development influence factor as: where represents the local development influence factor, represents the exponential function with the natural constant e as the base, represents the scaling factor, which can be set to 80; represents the global development influence factor; Obtain the single search rate based on the fitness corresponding to each parameter vector after initial search, and determine the average search rate in multiple training processes based on the single search rate as: where represents the average search rate, denotes the single search rate corresponding to the \(s\)th time being selected for local development or global development, and \(N\) denotes the number of times being selected for local development or global development during the previous training process. denotes the maximum value function. denotes the n fitness corresponding to the parameter vector after the \(n\)th initialization search when it is selected for local development or global development for the \(s\)th time. denotes the n fitness corresponding to the parameter vector after the \(n\)th initialization search when it is selected for local development or global development for the \((s - 1)\)th time; According to the average search speed, determine the local development threshold parameter and the global development threshold parameter as: wherein, denotes the local development threshold parameter corresponding to the n parameter vector after the \(n\)th initialization search, denotes the global development threshold parameter corresponding to the n parameter vector after the \(n\)th initialization search, denotes the average search rate for global development, and the average search rate for local development; According to the local development influence factor and the local development threshold parameter, determine the local development probability as: ; wherein, denotes the local development probability corresponding to the n parameter vector after the \(n\)th initialization search.

[0040] According to the local development probability, adopt a position transfer strategy for local development to determine the parameter vector after local development; According to the global development influence factor and the global development threshold parameter, determine the global development probability as: ; wherein, denotes the global development probability corresponding to the n parameter vector after the \(n\)th initialization search; According to the global development probability, adopt a chaotic global search strategy for global development to determine the parameter vector after global development; For any parameter vector after initialization search, if the parameter vector has not undergone global development and local development, directly use the original parameter vector after initialization search as the parameter vector after balanced search; If the parameter vector has been globally developed or locally developed, then use the parameter vector after global development or the parameter vector after local development as the parameter vector after balanced search; If the parameter vector has been globally developed and locally developed, then use the simulated annealing algorithm to select the parameter vector after global development or the parameter vector after local development as the parameter vector after balanced search.

[0041] Since the parameter vector after local development will not have a sharp deterioration in position, therefore, when the fitness of the parameter vector after global development is greater than the fitness of the parameter vector after local development, then use the parameter vector after global development as the parameter vector after balanced search; if the fitness of the parameter vector after global development is less than the fitness of the parameter vector after local development, then use the probability acceptance method of the simulated annealing algorithm to accept the parameter vector after global development as the parameter vector after balanced search, and when it is not accepted, then use the parameter vector after local development as the parameter vector after balanced search.

[0042] Through the above selection operation, the algorithm can have a greater global search ability in the early and middle stages, and gradually use the parameter vector after local development as the parameter vector after balanced search in the later stage of the algorithm, which indicates that the algorithm gradually turns to local search, but retains a certain global search ability.

[0043] In some possible implementation manners, according to the local development probability, use a position transfer strategy for local development to determine the parameter vector after local development, including: Randomly generate a first decision factor within the interval (0, 1), and determine whether the first decision factor is less than the local development probability. If so, use the position transfer strategy for local development to determine the parameter vector after local development, otherwise do not perform local development; Using the position transfer strategy for local development to determine the parameter vector after local development includes: For any parameter vector after initial search, determine the transfer probability corresponding to other parameter vectors according to the fitness corresponding to the parameter vector after initial search; Wherein, represents the m th parameter vector after initial search to the k th parameter vector after initial search transfer probability, m = 1, 2,..., NP, k = 1, 2,..., NP, and m is not equal tok ; represents the transfer intensity factor of the parameter vector after the t th training process and the k th initialization search, represents the transfer intensity factor of the parameter vector after the t th training process and the m th initialization search, represents the transfer intensity factor of the parameter vector after the t th training process and the h th initialization search, h i = 1, 2, …, NP, represents the attenuation factor between the intervals (0, 1), represents the t (i - 1)th training process and the m th initialization search of the parameter vector's transfer intensity factor, represents the finer parameter of the transfer intensity factor, represents the t th training process and the m th initialization search of the fitness of the parameter vector; , The calculation process of is similar to that of

[0044] Determine the target cooperation vector for the parameter vector after the initialization search according to the transfer probability corresponding to the other parameter vectors; The target cooperation vector for the m th parameter vector after the initialization search can be determined according to the transfer probability corresponding to the other parameter vectors, and then using the roulette wheel strategy; According to the target cooperation vector, perform local exploitation on the parameter vector after the initialization search, and determine the parameter vector after local exploitation as: where, represents the t th training process and the m th parameter vector after the initialization search, represents the m th parameter vector after local exploitation, represents the search step size, represents 's target cooperation vector, represents 's Euclidean distance from In the embodiment of the present invention, a position transfer strategy is adopted for local development, which can effectively enable the parameter vector to learn information from other positions, complete local information interaction, thereby realizing local search. At the same time, during the search process, there is a higher probability of searching for better positions, while retaining the possibility of searching for some cross positions. It can improve the global search ability in the middle and early stages of the algorithm and improve the search fineness in the later stage of the algorithm.

[0045] In some possible implementation manners, according to the global development probability, a chaotic global search strategy is adopted for global development to determine the parameter vector after global development, including: Randomly generate a second decision factor within the interval (0, 1), and determine whether the second decision factor is less than the global development probability. If so, adopt the chaotic global search strategy for global development to determine the parameter vector after global development; otherwise, do not perform global development. Adopting the chaotic global search strategy for global development to determine the parameter vector after global development includes: Obtain the chaotic mapping factor in the current training process as: Where represents the chaotic mapping factor in the t th training process, and is randomly generated from the interval (0, 1) at the start of training. represents the chaotic mapping factor in the t +1th training process. represents a constant term within the interval (0, 1) and is set to 0.7. According to the search upper limit vector and the search lower limit vector, determine the search center vector as: ; where represents the search center vector, represents the search upper limit vector, represents the search lower limit vector; The search upper limit vector is used to represent the vector composed of the upper limits of the hyperparameters to be trained, and the search lower limit vector represents the vector composed of the lower limits of the hyperparameters to be trained. It should be noted that each dimension of the vector pointed to in the embodiment of the present invention represents a fixed hyperparameter, which is convenient for encoding and decoding.

[0046] According to the chaotic mapping factor, combined with the sine function, perform global development on the parameter vector after initialization search to determine the parameter vector after global development as: Where represents the t th training process and the gThe parameter vector after an initial search indicating the g parameter vector after the th global exploration, where represents a random number between (0, ), and represents the Euclidean distance between and .

[0047] In the embodiments of the present invention, a chaotic global search strategy is adopted for global exploration, which can provide a powerful global search ability for the algorithm and assist the algorithm to jump out of local optimal solutions.

[0048] Through the above training process, the algorithm provided by the embodiments of the present invention can provide a relatively powerful global search ability in the middle and early stages of training. When the algorithm reaches the later stage, it provides a relatively powerful local search ability, but retains the global search ability, enabling the algorithm to still perform global search, improving the training effect of the algorithm, and ultimately ensuring the accuracy of speech recognition.

[0049] In some possible implementation manners, feature extraction is performed on the denoised voice command data to obtain voice features, including: extracting MFCC (Mel-Frequency Cepstral Coefficients) of the denoised voice command data to obtain voice features.

[0050] It should be noted that the above MFCC is only a preferred example provided by the embodiments of the present invention, and any time-domain feature or frequency-domain feature of voice features can also be used as voice features to achieve a similar effect.

[0051] In some possible implementation manners, based on the word vector and the driving state information, a reinforcement learning algorithm is used to determine driving control parameters, including: collecting the driving state information; where the state information includes position, attitude, and speed; constructing the word vector and the driving state information into a current state space, and using the current state space as the input of the reinforcement learning algorithm to determine the selected action and obtain driving control parameters; Optionally, the reinforcement learning algorithm can be set to a DQN (Deep Q-Network) algorithm.

[0052] Among them, the action is selected from the action space, each action represents a driving control parameter, and the action space should include all driving control parameters.

[0053] After controlling the overhead crane according to the overhead crane control parameters, if no new voice command is generated, the word vectors in the new state space are all set to 0.

[0054] The overhead crane control parameters may include the forward rotation number of turns parameter and the reverse rotation number of turns parameter of the overhead crane displacement motor, and may include the forward rotation number of turns parameter and the reverse rotation number of turns parameter of the spreader motor. Thus, the overhead crane can be controlled according to the overhead crane control parameters.

[0055] Optionally, a pressure sensor may also be provided on the hook, and then the running speed of the overhead crane is adjusted according to the pressure sensing data. For example, the speed can be appropriately increased during no-load or light-load to improve work efficiency, while the speed needs to be reduced when approaching the target position or lifting heavy objects to achieve stable docking and precise alignment.

[0056] After performing an action, an experience pool should also be constructed, and then the network parameters of the policy network in the reinforcement learning algorithm are updated according to the experience pool, so that the policy network can perform more accurate actions.

[0057] An overhead crane control system based on speech recognition provided by the present invention collects the voice command data of the operator through a microphone array, and after denoising the voice command data, the denoised voice command data is obtained. Then, a deep learning algorithm is used to recognize the denoised voice command data to determine the text command data, and the text command data is parsed to determine the overhead crane control keywords, and the overhead crane control keywords are converted into word vectors. Finally, based on the word vectors and the state information of the overhead crane, a reinforcement learning algorithm is used to determine the overhead crane control parameters, and based on the overhead crane control parameters, the overhead crane is controlled to perform operations, which can achieve precise and rapid control of the overhead crane, is applicable to more scenarios, and improves the overhead crane control efficiency.

[0058] As Figure 2 shown, an embodiment of the present invention provides an overhead crane control method based on speech recognition, including: S201. Collect the voice command data of the operator through a microphone array, and after denoising the voice command data, obtain the denoised voice command data; S202. Use a deep learning algorithm to recognize the denoised voice command data to determine the text command data; S203. Parse the text command data to determine the overhead crane control keywords, and convert the overhead crane control keywords into word vectors; S204. Based on the word vectors and the state information of the overhead crane, use a reinforcement learning algorithm to determine the overhead crane control parameters; S205. Based on the overhead crane control parameters, control the overhead crane to perform operations to complete the overhead crane control based on speech recognition.

[0059] A driving control method based on speech recognition provided by an embodiment of the present invention can be applied to the above system technical solution, and its principle and beneficial effects are similar, so details will not be described here.

[0060] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0061] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1 a device for the functions specified in one block or multiple blocks.

[0062] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements in the process Figure 1 one process or multiple processes and / or blocks Figure 1 a device for the functions specified in one block or multiple blocks.

[0063] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 a device for the functions specified in one block or multiple blocks.

[0064] Those of ordinary skill in the art can understand that all or part of the steps in realizing the above facts and methods can be completed by instructing relevant hardware through a program. The involved program or the program mentioned can be stored in a computer-readable storage medium. When the program is executed, it includes the following steps: At this time, the corresponding method steps are introduced. The storage medium can be ROM / RAM, magnetic disk, optical disk, etc.

[0065] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A driving control system based on speech recognition, characterized in that Including: A voice acquisition module, which is used to collect the voice command data of the operator through a microphone array, and after denoising the voice command data, obtain the denoised voice command data; A voice recognition module, which is used to recognize the denoised voice command data by using a deep learning algorithm to determine the text command data; An instruction parsing module, which is used to parse the text command data to determine the driving control keywords and convert the driving control keywords into word vectors; A reinforcement learning module, which is used to determine the driving control parameters by using a reinforcement learning algorithm based on the word vector and the state information of the vehicle; A driving control module, which is used to control the vehicle to perform operations based on the driving control parameters and complete the driving control based on voice recognition.

2. The vehicle driving control system based on speech recognition according to claim 1, wherein The voice recognition module includes a voice recognition model deployment sub-module, a voice data feature extraction sub-module, and a voice feature recognition sub-module; The voice recognition model deployment sub-module is used to pre-deploy a voice recognition model for voice recognition by using a deep learning algorithm; The voice data feature extraction sub-module is used to extract features from the denoised voice command data to obtain voice features; The voice feature recognition sub-module is used to schedule the voice recognition model pre-deployed by the voice recognition model deployment sub-module to recognize the voice features to determine the text command data.

3. The vehicle driving control system based on speech recognition according to claim 2, wherein Pre-deploying a voice recognition model for voice recognition by using a deep learning algorithm includes: Initializing the hyperparameters of the deep learning algorithm by using a chaotic mapping sequence to determine a plurality of different parameter vectors; wherein, the parameter vector includes all or part of the hyperparameters to be optimized of the deep learning algorithm; Obtaining the loss function value of each parameter vector and determining the fitness corresponding to the parameter vector according to the loss function value; Determining the optimal parameter vector and the worst parameter vector according to the fitness corresponding to the parameter vector; Based on the optimal parameter vector, using a variable spiral position selection mechanism to perform an initial search on the parameter vector to determine the parameter vector after the initial search; For the parameter vector after the initial search, according to the worst parameter vector, using an adaptive probability response mechanism to perform a local and global balanced search on the parameter vector to determine the parameter vector after the balanced search; Judging whether the current training times are greater than or equal to the preset maximum training times. If so, re-determine the optimal parameter vector according to the parameter vector after the balanced search. Otherwise, return to the step of obtaining the fitness; According to the re-determined optimal parameter vector, determine the final hyperparameters of the deep learning algorithm, and deploy a voice recognition model for voice recognition according to the final hyperparameters of the deep learning algorithm.

4. The vehicle driving control system based on speech recognition according to claim 3, wherein Based on the optimal parameter vector, using a variable spiral position selection mechanism to perform an initial search on the parameter vector to determine the parameter vector after the initial search, including: Based on the current training times, using the hyperbolic tangent function to determine the adaptive inertia weight; Using the adaptive inertia weight to weight the optimal parameter vector to determine the weighted optimal parameter vector; For any parameter vector, obtain the first difference vector between the optimal parameter vector and the parameter vector; For any parameter vector, perform a position transformation on the parameter vector using a variable helix function to determine the parameter vector after the position transformation; According to the weighted optimal parameter vector, the first difference vector, and the parameter vector after the position transformation, determine the parameter vector after the initial search.

5. The vehicle driving control system based on speech recognition according to claim 4, wherein, For the parameter vector after the initial search, according to the worst parameter vector, use an adaptive probability response mechanism to perform a local and global balanced search on the parameter vector to determine the parameter vector after the balanced search, including: Obtain the fitness corresponding to each parameter vector after the initial search, and obtain the population dispersion degree according to the fitness corresponding to each parameter vector after the initial search; According to the population dispersion degree, obtain the local development influence factor and the global development influence factor; Obtain the single search rate according to the fitness corresponding to each parameter vector after the initial search, and determine the average search rate during multiple training processes according to the single search rate; According to the average search speed, determine the local development threshold parameter and the global development threshold parameter; According to the local development influence factor and the local development threshold parameter, determine the local development probability, and according to the local development probability, use a position transfer strategy to perform local development to determine the parameter vector after local development; According to the global development influence factor and the global development threshold parameter, determine the global development probability, and according to the global development probability, use a chaotic global search strategy to perform global development to determine the parameter vector after global development; For any parameter vector after the initial search, if the parameter vector has not undergone global development and local development, directly use the original parameter vector after the initial search as the parameter vector after the balanced search; If the parameter vector has undergone global development or local development, use the parameter vector after global development or the parameter vector after local development as the parameter vector after the balanced search; If the parameter vector has undergone both global development and local development, use an annealing simulation algorithm to select the parameter vector after global development or the parameter vector after local development as the parameter vector after the balanced search.

6. The vehicle driving control system based on speech recognition according to claim 5, characterized in that, According to the local development probability, use a position transfer strategy to perform local development to determine the parameter vector after local development, including: Randomly generate a first decision factor within the interval (0,1), and determine whether the first decision factor is less than the local development probability. If so, use a position transfer strategy to perform local development to determine the parameter vector after local development, otherwise do not perform local development; Using a position transfer strategy to perform local development to determine the parameter vector after local development, including: For any parameter vector after the initial search, determine the transfer probability corresponding to other parameter vectors according to the fitness corresponding to the parameter vector after the initial search; According to the transfer probability corresponding to the other parameter vectors, determine the target cooperation vector for the parameter vector after the initial search; Perform local exploitation on the parameter vector after the initial search according to the target cooperation vector, and determine the parameter vector after local exploitation.

7. The vehicle driving control system based on voice recognition according to claim 6, wherein According to the global exploitation probability, adopt a chaotic global search strategy for global exploitation to determine the parameter vector after global exploitation, including: Randomly generate a second decision factor within the interval (0, 1), and determine whether the second decision factor is less than the global exploitation probability. If so, adopt a chaotic global search strategy for global exploitation to determine the parameter vector after global exploitation; otherwise, do not perform global exploitation. Adopt a chaotic global search strategy for global exploitation to determine the parameter vector after global exploitation, including: Obtain the chaotic mapping factor in the current training process. Determine the search center vector according to the search upper limit vector and the search lower limit vector. Perform global exploitation on the parameter vector after the initial search by combining the sine function according to the chaotic mapping factor to determine the parameter vector after global exploitation.

8. The vehicle driving control system based on speech recognition according to claim 2, wherein Extract features from the denoised voice command data to obtain voice features, including: extracting the MFCC of the denoised voice command data to obtain voice features.

9. The vehicle driving control system based on speech recognition according to claim 1, wherein Based on the word vector and the state information of the vehicle, use a reinforcement learning algorithm to determine the vehicle control parameters, including: Collect the state information of the vehicle; wherein, the state information includes position, attitude, and speed. Construct the current state space with the word vector and the state information of the vehicle, and use the current state space as the input of the reinforcement learning algorithm to determine the selected action and obtain the vehicle control parameters. Among them, the action is selected from the action space, each action represents a vehicle control parameter, and the action space should include all vehicle control parameters.

10. A driving control method based on speech recognition, characterized in that, Include: Collect the voice command data of the operator through the microphone array, and after denoising the voice command data, obtain the denoised voice command data. Use a deep learning algorithm to identify the denoised voice command data to determine the text command data. Parse the text command data to determine the vehicle control keywords, and convert the vehicle control keywords into word vectors. Based on the word vector and the state information of the vehicle, use a reinforcement learning algorithm to determine the vehicle control parameters. Based on the vehicle control parameters, control the vehicle to perform operations to complete vehicle control based on voice recognition.

Citation Information

Patent Citations

  • Method and device for operating speech-controlled information system for vehicle

    CN104603871A

  • System for central control of one or more cranes

    CN108698804A

  • Operating device and loading crane having an operating device

    CN111295355A

  • Intelligent unmanned crown block voice control system

    CN114373457A

  • Working machine voice interaction method and system and working machine

    CN114400001A