A logistics management system and method based on a speech recognition model
The voice recognition model in logistics management systems addresses inefficiencies in traditional data entry methods by enabling real-time, accurate command interpretation and execution, reducing errors and labor costs through deep learning and multi-modal verification.
Patent Information
- Application Number
- CN202510486729.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The existing speech recognition technology has insufficient professional adaptability in logistics management, resulting in recognition bias and semantic understanding errors. The data preprocessing and signal cleaning performance under the challenges of environmental diversity are insufficient, which cannot meet the actual application needs.
The logistics management system based on the speech recognition model is adopted to realize logistics information collection, instruction analysis and scheduling decision-making through voice interaction, combined with deep neural networks and intention analysis models, and use the pre-constructed logistics scenario feature library and multimodal verification mechanism to accurately identify professional terms and generate executable instructions.
It improves the accuracy and flexibility of logistics scheduling, reduces the manual operation error rate, realizes real-time instruction transmission and efficient logistics management, and reduces labor costs and training costs.
Smart Images

Figure CN120031466B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of logistics management, and specifically relates to a logistics management system and method based on a speech recognition model. Background Art
[0002] In traditional logistics management, manual data entry, barcode scanning, and fixed terminals are mainly relied on to collect and process logistics information. These methods have the following deficiencies: strong dependence on manual operations; low response speed, etc.
[0003] In recent years, with the continuous maturity of deep learning and big data technologies, breakthroughs have been made in speech recognition, natural language understanding, and intelligent decision-making technologies. With the help of voice input, natural language processing, and feedback closed-loop mechanisms, zero-touch and real-time instruction interaction can be achieved, greatly reducing the cost of manual intervention and improving operation convenience and response speed.
[0004] Although existing speech recognition technologies and logistics management systems already have relatively mature applications, there are still the following deficiencies in their combination: insufficient professional adaptability. When dealing with professional terms, phrases, and operation instructions involved in logistics scenarios, recognition deviations and semantic understanding errors often occur, resulting in inaccurate issued instructions; environmental and diversity challenges. In the logistics site environment, factors such as noise, echo, and background interference are relatively complex. The performance of traditional speech acquisition devices and algorithms in data preprocessing and signal cleaning is insufficient to meet the actual application requirements. Based on the above problems, there is an urgent need for a logistics management system and method based on a speech recognition model. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the present invention proposes a logistics management system and method based on a speech recognition model, which realizes logistics information collection, instruction parsing, and scheduling decision-making through voice interaction, reduces the manual operation error rate, and improves the flexibility and real-time nature of logistics scheduling response.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A logistics management method based on a speech recognition model, comprising:
[0008] Collecting voice data of logistics operators and performing preprocessing;
[0009] Performing speech recognition in combination with a pre-constructed logistics scenario feature library, constructing a loose instruction candidate set, and generating an initial operation instruction, where the loose instruction candidate set is composed of operation verbs, key entities, and key parameters;
[0010] Invoking a preset intention analysis model to perform intention analysis on the initial operation instruction to generate an executable instruction;
[0011] Verify the feasibility of the executable instruction through a multimodal verification mechanism, and transmit the verified operation instruction to the logistics management system for execution.
[0012] Specifically, the speech recognition is combined with a pre-built logistics scenario feature library to generate an initial operation instruction, including:
[0013] Use a deep neural network to extract the speech features of the preprocessed speech data, convert the continuous speech into frame-level feature data, and obtain a speech feature sequence;
[0014] Use a decoder and a language model to decode the speech feature sequence to generate an initial text description;
[0015] Perform word segmentation and keyword extraction on the generated initial text, and preliminarily separate the operation verbs, key entities, and key parameters through natural language processing technology to construct a loose instruction candidate set;
[0016] Use the pre-built logistics scenario feature library to correct and match the loose instruction candidate set;
[0017] Convert the corrected and matched text into a standardized initial operation instruction and output it.
[0018] Specifically, the use of the pre-built logistics scenario feature library to correct and match the loose instruction candidate set includes:
[0019] Compare the loose instruction candidate set with a preset operation action word library to correct the recognition deviation caused by speech homophony, dialect, and accent;
[0020] Verify and complete the key information, check and correct the data format, where the key information includes the warehouse number and the goods category;
[0021] Further screen and confirm the operation context according to the semantic rules of the logistics scenario context.
[0022] Specifically, the invocation of a preset intention analysis model to perform intention analysis on the initial operation instruction to generate an executable instruction includes:
[0023] Obtain the context data related to the current logistics scenario, including: historical operation records, real-time on-site status, business rules and constraints, and user and weight information;
[0024] Combine the context data and use the preset intention analysis model to perform intention analysis on the initial operation instruction;
[0025] Generate a standardized executable instruction according to the intention analysis result and output it.
[0026] Specifically, when using the combined context data to perform intent analysis on the initial operation instruction, it includes:
[0027] Integrate the initial operation instruction with the obtained context information into a unified data input, and input it into the preset intent analysis model;
[0028] Identify the main requested operation intent in the initial operation instruction, classify it, and extract the key parameters in the initial operation instruction, including warehouse number, quantity of goods, target area, vehicle information, time node;
[0029] Output the intent analysis result, including intent category, key parameters, recognition confidence, and marked uncertain information.
[0030] Specifically, when generating a standardized executable instruction based on the intent analysis result and outputting it, it includes:
[0031] Based on the intent analysis result, complete and perform rule verification on the key parameters;
[0032] Assemble the completed and verified data into an executable instruction according to the internal instruction standard and convert it into machine language.
[0033] Specifically, when completing and performing rule verification on the key parameters according to the intent analysis result, it includes:
[0034] For the initial operation instruction with clear intent and complete key parameters, directly assemble it into a standardized format according to business rules;
[0035] For missing or ambiguous key parameters, use the context information and preset default rules to complete them;
[0036] Use the format rule and data verification mechanism to make all key parameters conform to the preset format.
[0037] A logistics management system based on a speech recognition model for implementing the described logistics management method based on a speech recognition model, including: a data collection module, an initial instruction module, an executable instruction module, and a verification and transmission module;
[0038] The data collection module is used to collect the voice data of logistics operators and perform preprocessing;
[0039] The initial instruction module is used to perform speech recognition in combination with a pre-constructed logistics scenario feature library, construct a loose instruction candidate set, and generate an initial operation instruction. The loose instruction candidate set consists of operation verbs, key entities, and key parameters;
[0040] The executable instruction module is used to call a preset intention analysis model to analyze the intention of the initial operation instruction and generate an executable instruction;
[0041] The verification and transmission module is used to verify the feasibility of the executable instruction through a multimodal verification mechanism and transmit the verified operation instruction to the logistics management system for execution.
[0042] Specifically, the executable instruction module includes: a context processing unit, an intention analysis unit, and an executable instruction generation unit;
[0043] The context processing unit is used to obtain context data related to the current logistics scenario, including: historical operation records, real-time on-site status, business rules and constraints, and user and weight information;
[0044] The intention analysis unit is used to analyze the intention of the initial operation instruction by using a preset intention analysis model in combination with the context data;
[0045] The executable instruction generation unit is used to generate a standardized executable instruction according to the intention analysis result and output it.
[0046] Compared with the prior art, the beneficial effects of the present invention are:
[0047] 1. The present invention proposes a logistics management method based on a speech recognition model. The speech recognition model customized for the logistics industry and the expanded professional vocabulary can accurately identify professional terms, warehouse numbers, train numbers, etc., reduce the probability of misrecognition, ensure the high accuracy of operation instructions, and use historical data, real-time sensor information, image and video monitoring, and business rules to perform multi-dimensional verification on the instructions, further complement key parameters and correct recognition errors, providing an accurate data basis for subsequent automatic scheduling.
[0048] 2. The present invention proposes a logistics management method based on a speech recognition model. Through real-time speech collection, signal preprocessing, rapid recognition, and instruction generation, all business links are linked in real time, realizing the instantaneous transmission and scheduling of logistics information, greatly improving the overall logistics efficiency, significantly reducing errors and delays caused by manual input, reducing repetitive labor, and thus reducing labor costs and training costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a flowchart of a logistics management method based on a speech recognition model provided by the present invention;
[0050] Figure 2 It is an architecture diagram of a logistics management system based on a speech recognition model provided by the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0051] The present application is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any form. It should be noted that, for those of ordinary skill in the art, several variations and improvements can also be made without departing from the concept of the present application. These all belong to the protection scope of the present application.
[0052] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0053] It should be noted that, if there is no conflict, the various features in the embodiments of the present application can be combined with each other and are all within the scope of protection of the present application. In addition, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in a different order than the module division in the device or the order in the flow chart. In addition, the "
[0054] The words "first", "second", "third", etc. do not limit the data and execution order, but only distinguish the same items or similar items with basically the same functions and effects.
[0055] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used in this specification and in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The term "and / or" used in this specification includes any and all combinations of one or more of the related listed items.
[0056] Example 1
[0057] See also Figure 1 , an embodiment provided by the present invention: a logistics management method based on a speech recognition model, comprising the following specific steps:
[0058] Step S1: dynamically adjust voice collection parameters through environmental perception, collect voice data of logistics operators, and perform preprocessing;
[0059] In actual application scenarios, voice collection equipment can cover a variety of types, including portable mobile terminals (such as smart phones, tablets), fixed voice collection terminals, and vehicle-mounted voice collection equipment. In order to ensure stable collection effects in different scenarios such as logistics transportation, warehouse management, and dispatching command centers, the equipment's anti-noise design, microphone sensitivity, and multi-channel recording capabilities need to be considered when selecting the model;
[0060] The preprocessing includes preliminary processing, denoising, and cleaning. The preliminary processing includes time-domain and frequency-domain preprocessing, which perform time-domain filtering and frequency-domain conversion on the sampled original signal, and use techniques such as frame segmentation and window functions to divide the continuous speech into continuous short-time processing units. For denoising, in a noisy environment, a multi-channel microphone array is used for collection, noise is located through spatial information, and the background noise is effectively suppressed using beamforming algorithms. For single-channel devices, the noise reduction strategy is dynamically adjusted by combining adaptive filtering techniques, so as to accurately separate the speech signal in different noise spectra. According to the collected speech features, methods such as spectral subtraction, Wiener filtering, and deep learning speech enhancement are used to perform noise reduction processing on the speech, reducing the interference of environmental noise on the speech recognition model. During the processing, the signal-to-noise ratio (SNR) of the speech and noise is detected in real time, and the noise reduction filter parameters are adjusted to ensure that the output audio signal has high clarity and stability. The cleaning includes: eliminating echoes, reducing the interference of environmental reflections, performing dynamic range compression processing on the signal after noise reduction processing, smoothing sudden loudness changes in the speech. At the same time, the gain of each frequency band is adjusted through an audio equalizer, making the energy of the speech signal more prominent in the key frequency bands, enabling accurate parsing of the speech recognition model.
[0061] Step S2: Perform speech recognition in combination with the pre-constructed logistics scenario feature library to generate an initial operation instruction;
[0062] The specific steps of step S2 are as follows:
[0063] Step S201: Use a deep neural network to extract the speech features of the preprocessed speech data, convert the continuous speech into frame-level feature data, and obtain a speech feature sequence;
[0064] Step S202: Use a decoder and a language model to decode the speech feature sequence to generate an initial text description;
[0065] Step S203: Perform word segmentation and keyword extraction on the generated initial text, and preliminarily separate the operation verbs, key entities, and key parameters through natural language processing techniques to construct a loose instruction candidate set;
[0066] Step S204: Use the pre-constructed logistics scenario feature library to correct and match the loose instruction candidate set;
[0067] In this embodiment, the logistics feature library includes the following sub-libraries: 1. Operation action word library, which includes commonly used operation verbs in the logistics system, such as: dispatch, outbound, inbound, transshipment, sorting, loading, unloading, etc., and includes corresponding variant words, synonyms, and colloquial expressions, such as "call a car" → "dispatch vehicle"; 2. Key entity word library, which includes common target object names and numbers in logistics business: warehouse number, site code, train number, cargo code, area name, personnel ID, etc.; 3. Attribute template and data format rule library, which is used to perform structural verification on entity information, such as the license plate number should be in a certain format, such as "粤B12345", and the warehouse number needs to match a specific numbering system, such as "B-3-7"; 4. Context semantic rule library, which describes the context elements that different operation types should have, for example: the "departure" operation needs to identify the train number, time, and destination at the same time, and the "transshipment" should contain the starting station, the target station, and the cargo identification;
[0068] Establish corresponding attribute libraries for key entities involved in the scenario, such as "warehouse number", "cargo category", "license plate number", "destination", "time node", etc., collect and standardize common serial numbers, numbering formats, area divisions and other information in the field, and ensure that key parameters can be accurately extracted in actual identification;
[0069] The specific steps of step S204 are:
[0070] Step S2041: Compare the loose command candidate set with the preset operation action vocabulary to correct the recognition deviation caused by homophones, dialects, accents, etc.;
[0071] For example, the recognition result is "mobilize car No. 3" → inferred as "car No. 3" through fuzzy comparison, and corrected based on acoustic similarity and semantic context; the user says "the time for dispatching is nine o'clock" → the language model context determines that "dispatching" should be "dispatching", and corrected to the standard action;
[0072] Step S2042: verify and complete key information, check and correct the data format, wherein the key information includes warehouse number, cargo category, etc.;
[0073] In this embodiment, it can be seen from the attribute template and data format rule base that, for example, warehouse number, goods category, etc. have a fixed format, and it is necessary to ensure that the data format is consistent with the pre-defined rules;
[0074] Step S2043: Further filter and confirm the operation context according to the logistics scenario context semantic rules.
[0075] In this embodiment, a certain action must contain other information, such as the "dispatch vehicle" operation must match the corresponding vehicle, time and destination information.
[0076] Step S205: Convert the corrected and matched text into a standardized initial operation instruction and output it.
[0077] Exemplarily, the user says, "Send vehicle No. 3 to the west warehouse." The initial draft of the speech recognition text is: "Send vehicle No. 3 to the west warehouse." Through vocabulary matching, it is recognized that "fachai" is a common misrecognition item, which is highly similar to "send a vehicle". Matching the feature library, it is automatically corrected to "send a vehicle";
[0078] By pre - constructing a logistics scenario feature library, accurate matching of industry - specific terms and operation processes is achieved, improving the accuracy and response speed of speech recognition. Using the rules in the feature library for automatic correction reduces the risk of misrecognition caused by environmental interference and accent differences. Finally, natural language instructions are automatically converted into a structured format, providing clear and accurate data support.
[0079] Step S3: Invoke a preset intention analysis model to analyze the intention of the initial operation instruction and generate an executable instruction;
[0080] The specific steps of Step S3 are as follows:
[0081] Step S301: Obtain context data related to the current logistics scenario, including: historical operation records, real - time on - site status, business rules and constraints, and user and weight information;
[0082] Step S302: Combine the context data and use a preset intention analysis model to analyze the intention of the initial operation instruction;
[0083] The specific steps of Step S302 are as follows:
[0084] Step S3021: Integrate the initial operation instruction and the obtained context information into a unified data input and input it into the preset intention analysis model;
[0085] In this embodiment, after fusing the context information and the initial instruction, historical data and on - site status are fully utilized to improve the depth of the model's semantic understanding. Through unified format processing, the consistency of data input is ensured, reducing parsing ambiguities caused by format differences and providing an accurate data basis for subsequent intention recognition;
[0086] Step S3022: Identify the main requested operation intention in the initial operation instruction, classify it, and extract key parameters in the initial operation instruction, including warehouse number, quantity of goods, target area, vehicle information, time node, etc.;
[0087] In this embodiment, the model conducts in-depth analysis on the input data. First, it identifies the main requests and their corresponding operation types in the initial operation instructions and classifies them. For example, operations such as "position adjustment", "warehouse out", and "unloading" are categorized. During the intention recognition process, the model simultaneously extracts the key parameters involved in the instructions, such as: Warehouse number: Confirm the specific warehouse involved in the instruction; Quantity of goods: Identify the quantity of goods or transportation volume indicated in the instruction; Target area: Extract the destination or operation area; Vehicle information: Determine whether specific vehicles are involved in the scheduling; Time node: Identify the time requirement for executing the operation or the scheduling moment;
[0088] The preset model is used to accurately classify the initial instructions to accurately capture the user's operation requirements;
[0089] Step S3033: Output the intention analysis result, including the intention category, key parameters, recognition confidence, and marks for existing uncertain information.
[0090] In this embodiment, after completing intention recognition and key parameter extraction, the intention analysis result is output. The output result includes: Intention category: Clearly define the operation type (such as scheduling, unloading, warehousing, etc.) for subsequent classification processing; Key parameters: Include warehouse number, quantity of goods, target area, vehicle information, time node, etc., forming a complete set of instruction parameters; Recognition confidence: Evaluate the confidence of each parameter and the overall instruction, providing an indicator to judge the accuracy of the instruction; Uncertain information mark: Conduct a preliminary verification on the recognized information, detect whether there is ambiguous or low-confidence content, and mark this information with uncertainty;
[0091] This greatly improves the automated scheduling efficiency of the logistics system under voice interaction. At the same time, through multi-level data verification and feedback closed-loop, the accuracy and reliability of the instructions are effectively guaranteed.
[0092] Step S303: Generate a standardized executable instruction based on the intention analysis result and output it.
[0093] The specific steps of Step S303 are as follows:
[0094] Step S3031: Complement and perform rule verification on the key parameters according to the intention analysis result;
[0095] Complementing and performing rule verification on the key parameters includes:
[0096] For the initial operation instructions with clear intentions and complete key parameters, directly assemble them into a standardized format according to business rules;
[0097] For missing or ambiguous key parameters, use context information and preset default rules to complement them;
[0098] For example, if the scheduling time is not specified, the current system time is used by default;
[0099] Utilize formatting rules and data verification mechanisms to ensure that all key parameters conform to the preset format.
[0100] Such as warehouse codes, license plate number formats, etc.; The completion rules and steps here are different from the verification and matching in step S2. In the initial voice text, the amount of information contained is relatively large. Matching verification only preliminarily extracts useful information, that is, the initial operation instructions.
[0101] Step S3032: Assemble the completed and verified data into an executable instruction according to the internal instruction standard and convert it into machine language.
[0102] Specifically, before generating the instruction, it is necessary to verify and check the business logic to see if there are any conflicts or irrationalities among the key parameters. For example, whether the warehouse status matches the scheduling requirements; For data with low confidence or missing data, automatically generate warning or prompt messages and pass these messages to the operator for confirmation or supplementation; If serious defects are detected, block the execution of the instruction and wait for further manual intervention;
[0103] The main fields of the executable instruction include: command_type: operation category, such as dispatch_unload, store_in, etc.; warehouse_id: the verified warehouse number; destination_zone: the target area or operation location; vehicle_id: relevant vehicle information if scheduling is involved; dispatch_time: execution time; additional_parameters: other important parameters extracted by the model; confidence: the confidence level of overall recognition and parsing.
[0104] Step S4: Verify the feasibility of the executable instruction through a multimodal verification mechanism and transmit the verified operation instruction to the logistics management system for execution.
[0105] In this embodiment, the multimodal verification mechanism verification utilizes the verification of multimodal data (including but not limited to text, voice, images, sensor data, real-time GPS positioning, historical data, and business rules) to evaluate the feasibility of the generated operation instruction in terms of key parameter integrity, logical consistency, and matching degree with the actual logistics scenario. The operation instruction that passes the verification is transmitted to the logistics management system through a safe and stable data channel, prompting the subsequent scheduling execution module to immediately respond and execute the instruction;
[0106] The specific operation process is as follows: By semantic understanding, compare whether the current instruction matches the on-site image and sensor data, so as to supplement or correct the uncertain parts in the instruction; Compare the instruction parameters with the real-time status data, and give a comprehensive confidence score of the verification result to ensure the accuracy and executability of the instruction;
[0107] For abnormal states that occur during the verification process, such as missing key parameters and resource scheduling conflicts, the system automatically triggers an early warning mechanism: If the abnormality is within the tolerable range, the system attaches an abnormality mark to the instruction to prompt subsequent manual review; If the abnormality is serious, the instruction issuance is blocked, and a feedback report is generated for the operator to intervene and handle.
[0108] Through step S4, the multi-modal verification mechanism realizes strict verification of executable instructions at multiple levels such as the integrity of key parameters, logical consistency, and actual scenario adaptability, ensuring that each instruction issued to the logistics management system has high confidence and high reliability. At the same time, on the basis of ensuring data security and timely response, it realizes abnormal risk monitoring and feedback closed-loop, providing a solid technical support for the automated and efficient operation of the logistics scheduling system.
[0109] Embodiment 2
[0110] Please refer to Figure 2 , another embodiment provided by the present invention: A logistics management system based on a speech recognition model, including: a data acquisition module, an initial instruction module, an executable instruction module, and a verification and transmission module;
[0111] The data acquisition module is used to collect the voice data of logistics operators and perform preprocessing;
[0112] The initial instruction module is used to perform speech recognition in combination with a pre-constructed logistics scenario feature library, construct a loose instruction candidate set, and generate an initial operation instruction. The loose instruction candidate set consists of operation verbs, key entities, and key parameters;
[0113] The executable instruction module is used to call a preset intention analysis model to perform intention analysis on the initial operation instruction and generate an executable instruction;
[0114] The verification and transmission module is used to verify the feasibility of the executable instruction through a multi-modal verification mechanism and transmit the verified operation instruction to the logistics management system for execution.
[0115] The executable instruction module includes: a context processing unit, an intention analysis unit, and an executable instruction generation unit;
[0116] The context processing unit is configured to obtain context data related to the current logistics scenario, including: historical operation records, real-time on-site status, business rules and constraints, and user and weight information;
[0117] The intent analysis unit is configured to perform intent analysis on the initial operation instruction by using a preset intent analysis model in combination with the context data;
[0118] The executable instruction generation unit is configured to generate a standardized executable instruction according to the intent analysis result and output it.
[0119] In addition, parts of the above technical solutions provided in the embodiments of the present application that are consistent with the implementation principles of the corresponding technical solutions in the prior art are not described in detail to avoid excessive elaboration.
[0120] As described above in the specific embodiments, the objectives, technical solutions, and beneficial effects of the present invention have been further described in detail. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A logistics management method based on a speech recognition model, characterized in that It includes: Collect the voice data of logistics operators and perform preprocessing. Perform speech recognition in combination with a pre-constructed logistics scenario feature library, construct a loose instruction candidate set, and generate an initial operation instruction, where the loose instruction candidate set consists of operation verbs, key entities, and key parameters. Call a preset intention analysis model to analyze the intention of the initial operation instruction and generate an executable instruction. Verify the feasibility of the executable instruction through a multimodal verification mechanism, and transmit the verified operation instruction to the logistics management system for execution. The step of performing speech recognition in combination with a pre-constructed logistics scenario feature library to generate an initial operation instruction includes: Use a deep neural network to extract the speech features of the preprocessed voice data, convert continuous speech into frame-level feature data, and obtain a speech feature sequence. Use a decoder and a language model to decode the speech feature sequence to generate an initial text description. Perform word segmentation and keyword extraction on the generated initial text, and preliminarily separate operation verbs, key entities, and key parameters through natural language processing technology to construct a loose instruction candidate set. Use the pre-constructed logistics scenario feature library to correct and match the loose instruction candidate set. Convert the corrected and matched text into a standardized initial operation instruction and output it. The step of using the pre-constructed logistics scenario feature library to correct and match the loose instruction candidate set includes: Compare the loose instruction candidate set with a preset operation action word library to correct the recognition deviation caused by homophony, dialect, and accent in speech. Verify and complete the key information, check and correct the data format, where the key information includes warehouse numbers and goods categories. Further screen and confirm the operation context according to the semantic rules of the logistics scenario context.
2. The logistics management method based on a speech recognition model according to claim 1, characterized in that, The step of calling a preset intention analysis model to analyze the intention of the initial operation instruction and generate an executable instruction includes: Obtain the context data related to the current logistics scenario, including historical operation records, real-time on-site status, business rules and constraints, and user and weight information. Combine the context data and use a preset intention analysis model to analyze the intention of the initial operation instruction. Generate a standardized executable instruction according to the intention analysis result and output it.
3. The logistics management method based on a speech recognition model according to claim 2, characterized in that, The step of combining the context data and using a preset intention analysis model to analyze the intention of the initial operation instruction includes: Integrate the initial operation instruction with the obtained context information into a unified data input and input it into the preset intention analysis model. Identify the main requested operation intention in the initial operation instruction, classify it, and extract the key parameters in the initial operation instruction, including warehouse numbers, quantity of goods, target areas, vehicle information, and time nodes. Output the intention analysis result, including intention category, key parameters, recognition confidence, and marked uncertain information.
4. A logistics management method based on a speech recognition model according to claim 3, characterized in that The step of generating a standardized executable instruction according to the intention analysis result and outputting it includes: Complete and verify the rules of the key parameters according to the intention analysis result. Assemble the completed and verified data into an executable instruction according to the internal instruction standard and convert it into machine language.
5. A logistics management method based on a speech recognition model according to claim 4, characterized in that Based on the intention analysis result, complete and perform rule verification on the key parameters, including: For the initial operation instruction with a clear intention and complete key parameters, directly assemble it into a standardized format according to the business rules; For missing or ambiguous key parameters, use the context information and preset default rules to complete them; Use the format rules and data verification mechanism to make all key parameters conform to the preset format.
6. A logistics management system based on a speech recognition model for implementing a logistics management method based on a speech recognition model as described in any one of claims 1-5, characterized in that, Including: Data acquisition module, initial instruction module, executable instruction module, and verification and transmission module; The data acquisition module is used to collect the voice data of logistics operators and perform preprocessing; The initial instruction module is used to perform voice recognition in combination with the pre-built logistics scenario feature library, construct a loose instruction candidate set, and generate an initial operation instruction. The loose instruction candidate set consists of operation verbs, key entities, and key parameters; The executable instruction module is used to call the preset intention analysis model to perform intention analysis on the initial operation instruction and generate an executable instruction; The verification and transmission module is used to verify the feasibility of the executable instruction through a multimodal verification mechanism and transmit the verified operation instruction to the logistics management system for execution.
7. The logistics management system based on a speech recognition model according to claim 6, wherein, The executable instruction module includes: a context processing unit, an intention analysis unit, and an executable instruction generation unit; The context processing unit is used to obtain context data related to the current logistics scenario, including: historical operation records, real-time on-site status, business rules and constraints, and user and weight information; The intention analysis unit is used to perform intention analysis on the initial operation instruction by using the preset intention analysis model in combination with the context data; The executable instruction generation unit is used to generate a standardized executable instruction according to the intention analysis result and output it.
Citation Information
Patent Citations
Logistics management method and system based on speech recognition
CN108537480A
Intelligent interaction method and device suitable for voice information
CN113687719A