Agricultural product wholesale order automatic generation method and device, equipment and medium

By using microphone arrays and large language model technology, the voice signals of buyers and sellers in agricultural product wholesale telephone communications are separated, semantic understanding and entity information extraction are performed, and orders are generated and displayed. This solves the problems of low order generation efficiency and accuracy caused by mixed sound sources and semantic complexity, and achieves efficient and accurate order generation and real-time visual interaction.

CN121544346AActive Publication Date: 2026-02-17CHENGDU GAUSS ZHIDA INFORMATION TECHNOLOGY CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511729195.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-24
Publication Date
2026-02-17
Estimated Expiration
2045-11-24

AI Technical Summary

Technical Problem

Existing technologies suffer from low order generation efficiency, inaccurate information, and poor user experience in agricultural product wholesale telephone communication scenarios due to mixed sound sources, complex semantics, and lack of real-time interaction.

Method used

The system uses a microphone array to collect real-time sound signals from the site. It combines empirical mode decomposition, sound source localization, and voiceprint matching technologies to separate the voice signals of both buyers and sellers. It then uses a large language model to perform semantic understanding and entity information extraction, generating and displaying agricultural product wholesale orders in real time.

Benefits of technology

It enables efficient and accurate order generation in the context of telephone communication for agricultural product wholesale, improves order generation efficiency and accuracy, provides real-time visual interactive confirmation, frees up manpower, and enhances user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121544346A_ABST
    Figure CN121544346A_ABST
Patent Text Reader

Abstract

The invention discloses an agricultural product wholesale order automatic generation method and device, equipment and a medium, and relates to the technical field of order management. The method comprises the following steps: firstly, receiving a plurality of field sound signals collected by a sound pickup array on an agricultural product selling communication field in real time, and carrying out voice signal extraction processing by applying an empirical mode decomposition technology, a sound source positioning technology and a voiceprint matching technology in real time so as to obtain speech signals of a seller and a buyer; then, the extraction result is integrated into an agricultural product selling communication dialogue text data stream in real time, the text data stream is imported into a large language model in real time to carry out semantic understanding and entity information extraction processing, then a matched agricultural product wholesale order template is retrieved according to the extraction result, and content filling and dynamic updating are carried out on the template; and generating a new agricultural product wholesale order, and finally pushing the order to the seller terminal equipment in real time and outputting and displaying the order in real time, so that the order generation efficiency can be improved, the order information is ensured to be generated correctly, and the seller experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of order management, and particularly relates to an agricultural product wholesale order automatic generation method, device, equipment and medium. BACKGROUND

[0002] Agricultural product wholesale transactions have the characteristics of high transaction frequency, strong communication dependence and structured order information. At present, most transactions are still completed through telephone communication. The seller needs to manually record product name, specification, quantity and price information while answering the buyer's phone, and then enter the computer system or write the order. This way is inefficient, prone to errors, and cannot confirm the order details in real time during communication, which seriously affects transaction efficiency and accuracy.

[0003] To improve the level of transaction automation, some order generation schemes based on speech recognition have appeared in the prior art. For example, the existing patent "CN119831707A+ Agricultural product wholesale market voice automatic order generation system based on artificial intelligence" discloses a system for automatically generating orders using speech recognition technology, which collects voice signals, performs speech recognition and generates orders, to some extent, reducing the manual input link; the prior art such as "CN106056424A+ Agricultural product voice quick ordering system" proposes to collect the voice of the purchaser or seller through a voice receiving and recognizing module, and convert it into text data to complete the order; "CN109191269A+ A commodity transaction system based on voice interaction" aims to build a more natural shopping environment through voice interaction, but its recognition ability is not perfect; "CN203666729A+ Sales order acquisition method based on voice recognition" focuses on the telephone sales scenario, and fills in the address information in the customer's voice to the order template to improve the address recording efficiency.

[0004] Although the existing technology is aware of the application potential of voice technology in order generation to some extent, it still has obvious limitations in dealing with complex agricultural product wholesale telephone communication scenarios, which are specifically manifested as follows: (1) The sound source separation ability is insufficient, and the conversation between the two parties cannot be accurately distinguished, that is, most of the existing schemes rely on single voice signal collection and recognition, and lack the ability to effectively distinguish and separate the buyer's voice from the seller's voice in mixed voice signals (especially the mixed voice of the buyer's voice played through the telephone speaker and the seller's local voice); for example, the aforementioned prior art "CN119831707A" does not involve in-depth sound source separation processing of mixed voice signals, which makes the system unable to accurately determine which party proposes an order information (such as product quantity and specification), resulting in a blurred responsibility subject of the generated order, which is easy to cause subsequent disputes; (2) semantic understanding depth is not enough, order template matching and filling intelligence degree is low, that is, the existing technology focuses more on converting voice into text or performing simple keyword matching, lacks the ability of deep semantic understanding and entity extraction based on context, and cannot accurately capture the complex entity relationship in agricultural product transaction (such as multiple attributes such as category, grade and variety in "first-class Fuji apple"), and it is more difficult to dynamically drive the retrieval, adaptive filling and updating of the order template according to the dynamic dialogue content in real time; (3) the system lacks real-time interaction and verification feedback mechanism, that is, the existing scheme mostly focuses on order generation, but ignores real-time interaction and verification in the communication process; the seller cannot see the order draft automatically generated by the system during the call, so he cannot confirm and correct it in time according to the order content and the buyer; and there is no real-time checking and warning mechanism for missing mandatory items, optional items and logical errors (such as quantity unit mismatch), so the order accuracy still highly depends on manual post-check, which cannot fundamentally liberate manpower and improve efficiency.

[0005] In summary, the existing technology cannot effectively solve the problems of low order generation efficiency, inaccurate information and poor experience caused by mixed sound sources, complex semantics and lack of real-time interaction in the agricultural product wholesale telephone communication scene. Therefore, there is an urgent need for an automatic order generation scheme that can separate the voices of the two parties in real time, deeply understand the transaction semantics and support real-time visual confirmation. SUMMARY

[0006] The purpose of the present application is to provide an agricultural product wholesale order automatic generation method, device, computer equipment, computer readable storage medium and computer program product, to solve the problems of low order generation efficiency, inaccurate information and poor experience caused by mixed sound sources, complex semantics and lack of real-time interaction in the existing order generation technical scheme in the agricultural product wholesale telephone communication scene.

[0007] In order to achieve the above purpose, the present application adopts the following technical scheme: In a first aspect, an agricultural product wholesale order automatic generation method is provided, comprising: receiving a plurality of field sound signals collected in real time by a microphone array from an agricultural product selling communication field, wherein the microphone array comprises a plurality of microphones corresponding one-to-one to the plurality of field sound signals; applying empirical mode decomposition technology, sound source positioning technology and voiceprint matching technology to the plurality of field sound signals for real-time speech signal extraction processing to obtain seller speech signal and buyer speech signal; respectively performing real-time voice-to-text processing on the seller speech signal and the buyer speech signal to obtain seller speech text data stream and buyer speech text data stream; According to the seller speech text data stream and the buyer speech text data stream, a real-time agricultural product selling communication dialogue text data stream is formed; The agricultural product selling communication dialogue text data stream is introduced into a first large language model in real time for semantic understanding and entity information extraction processing, and a semantic understanding and entity information extraction result is obtained; According to the semantic understanding and entity information extraction result, a matched agricultural product wholesale order template is automatically retrieved from an order template library, and the agricultural product wholesale order template is filled in and dynamically updated in content to generate a new agricultural product wholesale order; The new agricultural product wholesale order is pushed to a seller terminal device in real time and output for display.

[0008] Based on the above invention, a new scheme of automatic order generation is provided, which can separate the speech of both parties in real time, deeply understand the transaction semantics, and support real-time visual confirmation. That is, first, multiple on-site sound signals collected by an array of sound pickers in real time during agricultural product selling communication are received, and empirical mode decomposition technology, sound source positioning technology and voiceprint matching technology are applied in real time for speech signal extraction processing to obtain seller and buyer speech signals. Then, the extraction results are integrated into an agricultural product selling communication dialogue text data stream in real time, and a large language model is introduced in real time for semantic understanding and entity information extraction processing. Then, a matched agricultural product wholesale order template is retrieved according to the extraction result, and the template is filled in and dynamically updated in content to generate a new agricultural product wholesale order. Finally, the order is pushed to a seller terminal device in real time and output for display. In this way, even in the context of agricultural product wholesale telephone communication, the generation efficiency of agricultural product wholesale orders can be improved, the correctness of order information generation can be ensured, the experience of sellers can be improved, and the scheme is convenient for practical application and promotion.

[0009] In one possible design, the empirical mode decomposition technology, the sound source positioning technology and the voiceprint matching technology are applied in real time to the multiple on-site sound signals for speech signal extraction processing to obtain seller speech signals and buyer speech signals, including: For each on-site sound signal in the multiple on-site sound signals, the corresponding signal is subjected to empirical mode decomposition processing in real time to obtain a plurality of intrinsic mode function components, and a plurality of component center frequency points corresponding to the plurality of intrinsic mode function components are determined, wherein the plurality of component center frequency points correspond one-to-one to the plurality of intrinsic mode function components; According to the plurality of component center frequency points of the each on-site sound signal, at least one adjacent frequency point group is found, wherein the adjacent frequency point group contains at least three component center frequency points corresponding to different speech signals, and the maximum frequency difference of the at least three component center frequency points is less than or equal to a preset frequency threshold; determining, for each of the at least one adjacent frequency point group, a corresponding sound source position according to at least three inherent modal function components corresponding to the at least three component center frequencies respectively and known arrangement positions of at least three microphones corresponding to the at least three inherent modal function components respectively; performing cluster analysis on the sound source positions of the adjacent frequency point groups to determine at least two sound source centers; determining, for each of the at least two sound source centers, all component center frequencies corresponding to any one of the plurality of field sound signals from all adjacent frequency point groups to which the sound source position belongs, and reconstructing an independent sound signal from the corresponding center according to all inherent modal function components corresponding to the all component center frequencies; respectively performing voiceprint matching processing on the independent sound signals of the sound source centers, and if the voiceprint of the independent sound signal of a certain sound source center is found to be the most matched with the voiceprint of the target salesperson, taking the independent sound signal of the certain sound source center as a seller speaking voice signal and taking the independent sound signal of another sound source center having a context dialogue relationship with the seller speaking voice signal as a buyer speaking voice signal.

[0010] In one possible design, according to the plurality of component center frequencies of the field sound signals, at least one adjacent frequency point group is found, including: sequentially examining the plurality of component center frequencies of the field sound signals along the frequency from small to large direction in the frequency domain, and if a certain component center frequency of a certain field sound signal is found to be located at a current frequency point and the current frequency point is not in a built frequency domain window, a frequency domain window with the current frequency point as a starting frequency point and a frequency domain width equal to a preset frequency threshold is newly built; For each of the built frequency domain windows, it is determined whether the total number of the plurality of component center frequencies corresponding to different sound signals in the corresponding window is greater than or equal to 3, and if yes, the plurality of component center frequencies are included in an adjacent frequency point group, wherein the adjacent frequency point group contains at least three component center frequencies corresponding to different sound signals, and the maximum frequency difference between the at least three component center frequencies is less than or equal to the preset frequency threshold.

[0011] In one possible design, for each of the at least one adjacent frequency point group, a corresponding sound source position is determined according to at least three inherent modal function components corresponding to the at least three component center frequencies respectively and known arrangement positions of at least three microphones corresponding to the at least three inherent modal function components respectively, including: For a certain adjacent frequency group in the at least one adjacent frequency group, determine at least three intrinsic mode function components that correspond one-to-one with the corresponding center frequency points of the at least three components; For each pair of intrinsic mode function components in the at least three intrinsic mode function components, the corresponding signal propagation time difference is calculated based on the corresponding two intrinsic mode function components; Based on the known placement positions of at least three microphones that correspond one-to-one with the at least three intrinsic mode function components and the propagation time difference values ​​of each pair of intrinsic mode function components, the sound source location corresponding to a certain adjacent frequency point group is calculated using a time difference localization algorithm.

[0012] In one possible design, an independent sound signal from another sound source center that has a contextual dialogue relationship with the seller's spoken voice signal is used as the buyer's spoken voice signal, including: The system sequentially traverses all other sound source centers among the at least two sound source centers in order of distance from the aforementioned sound source center: First, it performs speech-to-text processing on the seller's speech voice signal and the independent voice signals of the currently traversed other sound source centers to obtain the seller's speech text data stream and the other speech text data streams. Then, it imports the seller's speech text data stream and the other speech text data streams into the second language model to determine whether there is a contextual dialogue relationship. If there is, the independent voice signal of the currently traversed other sound source center is used as the buyer's speech voice signal; otherwise, it traverses the next other sound source center.

[0013] In one possible design, the semantic understanding and entity information extraction results include agricultural product name entity information, agricultural product specification entity extraction information, agricultural product quantity entity information, agricultural product price entity information, agricultural product transaction time entity information and / or agricultural product transaction location entity information; Based on the semantic understanding and entity information extraction results, matching agricultural product wholesale order templates are automatically retrieved from the order template library, including: Based on the agricultural product name entity information in the semantic understanding and entity information extraction results, agricultural product wholesale order templates that match the product category to which the agricultural product name entity information belongs are automatically retrieved from the order template library.

[0014] In one possible design, before pushing the new agricultural product wholesale order to the target salesperson's terminal device in real time, the method further includes: Check whether the required fields in the new agricultural product wholesale order have been filled in. If not, display the required field in the new agricultural product wholesale order with a first warning. And / or, check whether the optional fields in the new agricultural product wholesale order have been filled in; if not, present the optional field in the new agricultural product wholesale order in a prompt manner. And / or, check whether there are logical errors in the filled items in the new agricultural product wholesale order; if so, present the filled items in the new agricultural product wholesale order in a second warning manner.

[0015] Secondly, an automatic agricultural product wholesale order generation device is provided, which includes a sound signal receiving unit, a speech signal extraction unit, a speech-to-text conversion unit, a dialogue text forming unit, a semantic entity extraction unit, a wholesale order generation unit, and an order push display unit that are connected in sequence. The sound signal receiving unit is used to receive multiple on-site sound signals collected in real time by the microphone array at the agricultural product sales communication site, wherein the microphone array includes multiple microphones that correspond one-to-one with the multiple on-site sound signals. The speech signal extraction unit is used to apply empirical mode decomposition technology, sound source localization technology and voiceprint matching technology in real time to extract speech signals from the multiple on-site sound signals to obtain the seller's speech speech signal and the buyer's speech speech signal. The speech-to-text conversion unit is used to process the seller's speech voice signal and the buyer's speech voice signal into speech-to-text in real time to obtain the seller's speech text data stream and the buyer's speech text data stream. The dialogue text forming unit is used to form a real-time dialogue text data stream for agricultural product sales communication based on the seller's speech text data stream and the buyer's speech text data stream. The semantic entity extraction unit is used to import the agricultural product sales communication dialogue text data stream into the first language model in real time for semantic understanding and entity information extraction processing, and obtain semantic understanding and entity information extraction results. The wholesale order generation unit is used to automatically retrieve a matching agricultural product wholesale order template from the order template library based on the semantic understanding and entity information extraction results, and to fill in and dynamically update the content of the agricultural product wholesale order template to generate a new agricultural product wholesale order. The order push and display unit is used to push the new agricultural product wholesale order to the seller's terminal device in real time and display it in real time.

[0016] Thirdly, the present invention provides a computer device comprising a storage module, a processing module, and a transceiver module connected in sequence for communication, wherein the storage module is used to store a computer program, the transceiver module is used to send and receive messages, and the processing module is used to read the computer program and execute the automatic generation method for agricultural product wholesale orders as described in the first aspect or any possible design in the first aspect.

[0017] Fourthly, the present invention provides a computer-readable storage medium storing instructions that, when executed on a computer, perform the method for automatically generating agricultural product wholesale orders as described in the first aspect or any possible design within the first aspect.

[0018] Fifthly, the present invention provides a computer program product, including a computer program or instructions, which, when executed by a computer, implement the method for automatically generating wholesale agricultural product orders as described in the first aspect or any possible design in the first aspect.

[0019] The beneficial effects of the above scheme are: (1) This invention creatively provides a new automated order generation solution that can separate the voices of both parties in real time, deeply understand the semantics of transactions, and support real-time visual confirmation. First, it receives multiple on-site sound signals collected in real time by a microphone array from the communication scene of agricultural product sales. Then, it applies empirical mode decomposition technology, sound source localization technology, and voiceprint matching technology in real time to extract and process the voice signals of the seller and the buyer. Then, it integrates the extraction results into the text data stream of the agricultural product sales communication dialogue in real time and imports it into a large language model in real time for semantic understanding and entity information extraction. Then, it retrieves the matching agricultural product wholesale order template according to the extraction results, fills in the content of the template and updates it dynamically to generate a new agricultural product wholesale order. Finally, it pushes the order to the seller's terminal device in real time and outputs and displays it in real time. In this way, even in the scenario of agricultural product wholesale telephone communication, it can improve the generation efficiency of agricultural product wholesale orders and ensure that the order information is generated correctly, thereby improving the seller's experience. (2) It can fundamentally improve the efficiency and accuracy of order generation, that is, it realizes the full automation of the order generation process, completely liberating sellers from the high-intensity mental and physical labor of "listening-memorizing-manualizing"; specifically, through automatic collection, separation, identification, understanding and filling, communication and order generation are changed from "serial" to "parallel", and an accurate draft order is obtained after the communication ends, which improves efficiency several times and completely eliminates order errors caused by human mishearing, misremembering or omission. (3) It can overcome the challenges of sound source separation and identity recognition in complex environments. In a complex sound field that mixes telephone speaker voice and ambient noise, it can accurately and robustly separate and identify the voices of the buyer and seller. Specifically, through empirical mode decomposition, it can adaptively process non-stationary signals and effectively separate sound sources with different characteristics (clear electronic voice vs. rich natural voice) and noise in the frequency domain, providing a clean signal basis for subsequent positioning. Through sound source localization and clustering, and based on the "adjacent frequency point group" and TDOA algorithm, it can accurately locate the physical sound source and cluster the main sound source center, effectively resisting the interference of environmental echo and random noise. Through dual verification of voiceprint and context, combined with the pre-registered seller voiceprint and LLM's logical judgment of the dialogue content, it provides dual protection for the accurate determination of the identities of the buyer and seller, solving the misjudgment that may occur by relying solely on location or voiceprint as a single factor. (4) It can achieve deep semantic understanding and intelligent template adaptation, that is, it can not only "hear" the text, but also "understand" the transaction intent and drive the business system to make intelligent decisions and adapt; Specifically, it can extract the entity through the large language model and use the powerful semantic understanding capability of LLM to accurately extract the structured transaction entities (product name, specifications, quantity, price, etc.) scattered in the dialogue flow, and understand complex language phenomena such as reference, omission and correction; and through dynamic template retrieval, it can automatically match the most professional order template according to the identified product type (such as meat templates automatically including the "quarantine certificate number" field), realizing the intelligent upgrade from general to professional, and improving the standardization and professionalism of orders; (5) It can create a real-time, transparent and trustworthy interactive confirmation experience, that is, the traditional "black box" back-end processing process is transformed into a transparent collaborative process that is visible, error-correctable and confirmable; specifically, the order is presented in real time, and the seller can view the automatically generated order while communicating, so that they can confirm with the purchaser in real time based on the visualized content ("You want 100 catties of first-grade apples, right?"), which greatly enhances the accuracy and trust of communication; and through multi-level intelligent alerts, it can realize the revolutionary user experience of "chatting and modifying, and confirming as soon as it is generated", which is convenient for practical application and promotion. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating the method for automatically generating wholesale orders for agricultural products provided in this application embodiment.

[0022] Figure 2 An example diagram showing the frequency domain window establishment result provided in the embodiments of this application.

[0023] Figure 3 A schematic diagram of the structure of the automatic agricultural product wholesale order generation device provided in the embodiments of this application.

[0024] Figure 4 A schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the present invention will be briefly introduced below in conjunction with the accompanying drawings and descriptions of the embodiments or the prior art. Obviously, the following description of the structure of the accompanying drawings is only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these embodiments without creative effort. It should be noted that the description of these embodiments is for the purpose of helping to understand the present invention, but does not constitute a limitation of the present invention.

[0026] It should be understood that although the terms "first" and "second", etc., may be used herein to describe various objects, these objects should not be limited by these terms. These terms are only used to distinguish one object from another. For example, the first object may be referred to as the second object, and similarly, the second object may be referred to as the first object, without departing from the scope of the exemplary embodiments of the invention.

[0027] It should be understood that the term "and / or" that may appear in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, B exists alone, or A and B exist simultaneously. Another example is A, B and / or C, which can mean that any one of A, B, and C or any combination thereof exists. The term " / and" that may appear in this document describes another relationship between related objects, indicating that two relationships can exist. For example, A / and B can mean: A exists alone or A and B exist simultaneously. In addition, the character " / " that may appear in this document generally indicates that the related objects before and after it are in an "or" relationship.

[0028] Example like Figure 1As shown, the method for automatically generating wholesale agricultural product orders provided in the first aspect of this embodiment can be executed, but is not limited to, by a computer device with certain computing resources and a communication connection to a microphone array, such as a server, a personal computer (PC, referring to a multi-purpose computer of a size, price, and performance suitable for personal use; desktop computers, laptops, mini-laptops, tablets, and ultrabooks are all personal computers), a smartphone, a personal digital assistant (PDA), or a wearable device. Figure 1 As shown, the method for automatically generating wholesale orders for agricultural products includes, but is not limited to, the following steps S1 to S7.

[0029] S1. Receive multiple on-site sound signals collected in real time from the on-site communication of agricultural product sales by a microphone array, wherein the microphone array includes multiple microphones that correspond one-to-one with the multiple on-site sound signals.

[0030] In step S1, the microphone array needs to be deployed in the agricultural product sales communication scene. This scene can be a wholesale telephone communication scenario, a direct wholesale communication scenario, or so on. The microphones are used to independently and in real-time acquire the corresponding on-site sound signals. These on-site sound signals are the result of a live mix of all sound signals generated in the complex wholesale telephone / direct communication scenario, including, but not limited to, the seller's voice signal, the buyer's voice signal, the third party's voice signal, and other noise signals. For example, there can be six microphones arranged on the six faces of a dodecahedron. Furthermore, since the on-site sound signals are transmitted via data transmission, both the on-site sound signals and subsequent signals are digital signals.

[0031] S2. Real-time application of empirical mode decomposition technology, sound source localization technology, and voiceprint matching technology to extract and process the multiple on-site sound signals to obtain the seller's speech voice signal and the buyer's speech voice signal.

[0032] In step S2, since the on-site sound signal contains the seller's speech signal, the buyer's speech signal, the third party's speech signal, and other noise signals, it is necessary to purposefully separate and extract the seller's speech signal and the buyer's speech signal for subsequent processing steps such as speech recognition, semantic understanding, and information extraction. Empirical Mode Decomposition (EMD) is a classic existing denoising technique for processing non-stationary / nonlinear signals. Specifically, it performs empirical mode decomposition on the noisy signal, calculates the Intrinsic Mode Functions (IMFs) of each order, and then reconstructs certain IMF components to achieve the effect of high-pass and / or low-pass filters. The sound source localization technology is used to determine the source of different sound signals, so as to identify the source of the seller's voice signal (i.e., the mouth of the target salesperson) and the source of the buyer's voice signal (i.e., the mouth of the on-site buyer or telephone speaker, etc.) by combining with voiceprint matching technology. Then, the empirical mode decomposition technology, the sound source localization technology, and the voiceprint matching technology can be applied to extract the seller's voice signal and the buyer's voice signal. In order to achieve accurate extraction of voice signals, preferably, the empirical mode decomposition technology, the sound source localization technology, and the voiceprint matching technology are applied in real time to extract voice signals from the multiple on-site sound signals to obtain the seller's voice signal and the buyer's voice signal, including but not limited to the following steps S21 to S26.

[0033] S21. For each of the multiple on-site sound signals, perform empirical mode decomposition processing on the corresponding signal in real time to obtain multiple intrinsic mode function components, and determine multiple component center frequency points based on the multiple intrinsic mode function components, wherein the multiple component center frequency points correspond one-to-one with the multiple intrinsic mode function components.

[0034] In step S21, since the empirical mode decomposition technique assumes that any complex sequence is formed by the superposition of multiple single-frequency signals, it can be decomposed into a combination of several intrinsic mode functions (IMFs). Assuming the on-site sound signal is... ( (representing a time variable), then The EMD decomposition formula is:

[0035] In the formula, This represents the total number of the plurality of intrinsic mode function components. Indicates less than or equal to positive integers, Indicates the first Each intrinsic mode function component (i.e., a single-frequency signal). This represents the residual signal. The specific process of the empirical mode decomposition is existing technology and will not be described in detail here. Furthermore, since the intrinsic mode function component is a single-frequency signal, the spectrum of the intrinsic mode function component can be obtained by conventional signal time-domain to frequency-domain conversion (e.g., Fourier transform), and then the center point of the spectrum can be selected as the corresponding component center frequency.

[0036] S22. Based on the multiple component center frequency points of each of the on-site sound signals, at least one adjacent frequency point group is found, wherein the adjacent frequency point group contains at least three component center frequency points corresponding to different sound signals, and the maximum frequency difference of the at least three component center frequency points is less than or equal to a preset frequency threshold.

[0037] In step S22, the technical approach for finding the adjacent frequency point group is as follows: sound signals from the same sound source but arriving at different locations (i.e., the locations of different microphones) have the same or similar frequencies. Therefore, it can be inferred that different intrinsic mode function components with the same or similar frequencies may originate from the same sound source. That is, one adjacent frequency point group can be assumed to correspond to one sound source. Specifically, based on the center frequencies of the multiple components of each on-site sound signal, at least one adjacent frequency point group is found, including but not limited to the following steps S221 to S222.

[0038] S221. In the frequency domain, examine the center frequency points of the multiple components of each field sound signal in sequence along the direction from small to large frequency. If it is found that the center frequency point of a component of a certain field sound signal is located at the current frequency point and the current frequency point is not within the established frequency domain window, then create a new frequency domain window with the current frequency point as the starting frequency point and the frequency domain width equal to the preset frequency threshold.

[0039] In step S221, the approach to establishing the frequency domain window is as follows: among the multiple component center frequencies of each on-site sound signal, in the frequency domain along the direction from smallest to largest frequency, the first component center frequency is used as the trigger starting point (i.e. Figure 2 The frequency domain window is defined by drawing a line backward from the red dot (in the diagram), and the window width is set. (Note that the center frequency of a component within the window cannot trigger the creation of a new window; i.e., window overlap must be avoided.) Based on the example of six live sound signals (each corresponding to one of the six microphones), the frequency domain window is established as follows: Figure 2 As shown. Furthermore, the preset frequency threshold can be pre-set based on historical experience.

[0040] S222. For each of the established frequency domain windows, determine whether the total number of multiple component center frequency points corresponding to different sound signals within the corresponding window is greater than or equal to 3. If so, include the multiple component center frequency points into an adjacent frequency point group, wherein the adjacent frequency point group contains at least three component center frequency points corresponding to different sound signals, and the maximum frequency difference of the at least three component center frequency points is less than or equal to the preset frequency threshold.

[0041] In step S222, an example based on six on-site sound signals (each corresponding to one of the six microphones) is given, such as... Figure 2 As shown, a group of adjacent frequency points containing component center frequency point 1A, component center frequency point 3A, and component center frequency point 6A can be obtained.

[0042] S23. For each adjacent frequency group in the at least one adjacent frequency group, determine the corresponding sound source location based on the known placement of at least three intrinsic mode function components corresponding to the corresponding center frequency points of the at least three components and the known placement of at least three microphones corresponding to the at least three intrinsic mode function components.

[0043] In step S23, a location algorithm based on TDOA (Time Difference of Arrival) can be used to locate the sound source. Specifically, for each adjacent frequency group in the at least one adjacent frequency group, the location of the corresponding sound source is determined according to the known placement of at least three intrinsic mode function components corresponding to the center frequency points of the corresponding at least three components and at least three microphones corresponding to the at least three intrinsic mode function components. This includes, but is not limited to, the following steps S231 to S233.

[0044] S231. For a certain adjacent frequency group in the at least one adjacent frequency group, determine at least three intrinsic mode function components that correspond one-to-one with the corresponding center frequency points of the at least three components.

[0045] S232. For each pair of intrinsic mode function components in the at least three intrinsic mode function components, calculate the corresponding signal propagation time difference value based on the corresponding two intrinsic mode function components.

[0046] In step S232, since the two intrinsic mode function components are single-point frequency signals with the same or adjacent frequencies, the signal propagation time difference can be directly obtained based on their phase difference. Alternatively, it can be statistically obtained based on their actual peak / trough conditions. Specifically, for each pair of intrinsic mode function components among the at least three intrinsic mode function components, the corresponding signal propagation time difference is calculated based on the two corresponding intrinsic mode function components, including but not limited to steps S2321 to S2323.

[0047] S2321. For each pair of intrinsic mode function components in the at least three intrinsic mode function components, find at least one adjacent peak / trough time group based on the corresponding two intrinsic mode function components, wherein the adjacent peak / trough time group contains two peak / trough times that correspond one-to-one with the two intrinsic mode function components, and the time difference between the two peak / trough times is less than or equal to a preset time threshold.

[0048] In step S2321, the method for finding adjacent peak / trough time groups can be derived conventionally by referring to the method for finding adjacent frequency point groups, and will not be repeated here.

[0049] S2322. For each adjacent peak / trough time group in the at least one adjacent peak / trough time group, calculate the corresponding time difference based on the corresponding two peak / trough times.

[0050] S2323. Calculate the average time difference of each adjacent peak / trough time group to obtain the signal propagation time difference value corresponding to the pair of intrinsic mode function components.

[0051] S233. Based on the known placement positions of at least three microphones corresponding one-to-one with the at least three intrinsic mode function components and the propagation time difference values ​​of each pair of intrinsic mode function components, the sound source position corresponding to a certain adjacent frequency point group is calculated using a time difference positioning algorithm.

[0052] In step S233, since the at least three intrinsic mode function components correspond to different sound signals, and these different sound signals originate from different microphones, at least three microphones corresponding one-to-one with the at least three intrinsic mode function components can be identified. Furthermore, the time difference positioning algorithm is an existing positioning algorithm, and its specific algorithm principle will not be elaborated here.

[0053] S24. Perform cluster analysis on the sound source locations of each adjacent frequency group to determine at least two sound source centers.

[0054] In step S24, since the sound source positions of each adjacent frequency group can be regarded as three-dimensional coordinate data, the three-dimensional coordinate data of the at least two sound source centers can be determined based on the three-dimensional coordinate data and clustering analysis algorithms such as K-means algorithm, thereby eliminating the positioning deviation caused by the aforementioned sound source positioning technology and ensuring the accuracy of the sound source positioning results.

[0055] S25. For each of the at least two sound source centers, determine all component center frequencies corresponding to any one of the multiple field sound signals from all adjacent frequency point groups whose sound source locations belong to the corresponding center, and reconstruct an independent sound signal from the corresponding center based on all intrinsic mode function components that correspond one-to-one with all component center frequencies.

[0056] In step S25, all the aforementioned intrinsic mode function components are the sum of the sound signals generated by an independent sound source, and therefore the independent sound signals of each sound source center can be conventionally reconstructed.

[0057] S26. Perform voiceprint matching processing on the independent sound signals of each sound source center respectively. If the voiceprint of the independent sound signal of a certain sound source center is found to be the best match with the voiceprint of the target salesperson, then the independent sound signal of that sound source center is taken as the seller's speech voice signal, and the independent sound signal of another sound source center that has a contextual dialogue relationship with the seller's speech voice signal is taken as the buyer's speech voice signal.

[0058] In step S26, the voiceprint features of the target salesperson need to be collected and stored in advance. Then, during voiceprint matching processing, the matching degree between the voiceprint features of the independent sound signals of each sound source center and the voiceprint features of the target salesperson is calculated conventionally. Finally, the independent sound signal of the sound source center with the highest matching degree is taken as the seller's spoken voice signal. Since the communication between the buyer and seller has the characteristics of a question-and-answer dialogue, the seller's spoken voice signal and the buyer's spoken voice signal will have a contextual dialogue relationship. Therefore, the buyer's spoken voice signal can be accurately determined based on the presence or absence of this contextual dialogue relationship. Furthermore, considering the close-range communication characteristics of buyers and sellers (even during telephone communication, the buyer can communicate closely with the seller through the speakerphone), to quickly and accurately determine the buyer's spoken voice signal, it is further preferred to use an independent sound signal from another sound source center that has a contextual dialogue relationship with the seller's spoken voice signal as the buyer's spoken voice signal. This includes, but is not limited to, traversing each of the other sound source centers among the at least two sound source centers in order of distance from the aforementioned sound source center: first, performing speech-to-text processing on the seller's spoken voice signal and the independent sound signals of the currently traversed other sound source centers, respectively, to obtain the seller's spoken text data stream and the other spoken text data streams; then, importing the seller's spoken text data stream and the other spoken text data streams into a second language model to determine whether a contextual dialogue relationship exists. If it exists, the independent sound signal of the currently traversed other sound source center is used as the buyer's spoken voice signal; otherwise, the next other sound source center is traversed. The specific process of the speech-to-text processing can be conventionally implemented based on existing speech recognition algorithms and will not be elaborated here. Since using LLM (Large Language Model) for named entity recognition and relation extraction is a standard application in the current field of NLP (Natural Language Processing), the second large language model with contextual dialogue relation extraction capability can be used by regular fine-tuning to identify whether there is a contextual dialogue relationship between the seller's speech text data stream and the other speech text data streams.

[0059] S3. Perform real-time speech-to-text processing on the seller's speech signal and the buyer's speech signal to obtain the seller's speech text data stream and the buyer's speech text data stream.

[0060] In step S3, the specific process of speech-to-text processing can also be conventionally implemented based on existing speech recognition algorithms, and will not be described in detail here.

[0061] S4. Based on the seller's speech text data stream and the buyer's speech text data stream, generate an agricultural product sales communication dialogue text data stream in real time.

[0062] In step S4, since the seller's speech text data stream contains at least one seller's speech text arranged in chronological order, and the buyer's speech text data stream contains at least one buyer's speech text arranged in chronological order, the at least one seller's speech text and the at least one buyer's speech text can be arranged in chronological order to form the agricultural product sales communication dialogue text data stream.

[0063] S5. Import the agricultural product sales communication dialogue text data stream into the first language model in real time for semantic understanding and entity information extraction processing to obtain the semantic understanding and entity information extraction results.

[0064] In step S5, the semantic understanding and entity information extraction results form the information basis for selecting the order template and filling in the template. Specifically, this includes, but is not limited to, entity information related to the name of agricultural products (e.g., "cabbage," "apple," or "ribs"), entity information related to the specifications of agricultural products (e.g., "grade 1," "bulk," or "frozen"), entity information related to the quantity of agricultural products (e.g., "100 jin" or "20 boxes"), entity information related to the price of agricultural products (e.g., "2.5 yuan / jin"), entity information related to the transaction time of agricultural products (e.g., "tomorrow morning" or "3 pm"), and / or entity information related to the transaction location of agricultural products (e.g., "warehouse number three"). Furthermore, the semantic understanding and entity information extraction processing can also be achieved by routinely fine-tuning the first large language model, which possesses semantic understanding and entity information extraction capabilities.

[0065] S6. Based on the semantic understanding and entity information extraction results, automatically retrieve the matching agricultural product wholesale order template from the order template library, fill in the content of the agricultural product wholesale order template and dynamically update it to generate a new agricultural product wholesale order.

[0066] In step S6, the order template library pre-stores wholesale agricultural product order templates for different product categories (which include some required and optional fields). Specifically, based on the semantic understanding and entity information extraction results, matching wholesale agricultural product order templates are automatically retrieved from the order template library, including but not limited to: based on the agricultural product name entity information in the semantic understanding and entity information extraction results, automatically retrieving wholesale agricultural product order templates from the order template library that match the product category to which the agricultural product name entity information belongs. For example, when the entity information of the agricultural product name is "ribs", the "meat order template" can be automatically retrieved. This template may, but is not limited to, have pre-set fields such as "quarantine certificate number" and "slaughter date", as well as fields for quantity, price, transaction time and / or transaction location to be filled in. Furthermore, it can fill in the content of the corresponding fields based on the extracted information of agricultural product entity specifications, quantity, price, transaction time and / or transaction location from the semantic understanding and entity information extraction results. When there are updates to the extracted information of agricultural product entity specifications, quantity, price, transaction time and / or transaction location, the corresponding fields are dynamically updated (for example, when the buyer says "No, I want 50 jin, not 15 jin", the large language model can understand that this is a negation and correction of the previous information, and immediately update the agricultural product quantity entity information from "15" to "50", and thus update the quantity field from "15" to "50"), and finally generate a new agricultural product wholesale order. In addition, when the buyer says "same as last time", the first language model can also obtain other semantic understanding and entity information extraction results, so as to combine historical order data to automatically fill in specific products and quantities, and also generate new agricultural product wholesale orders.

[0067] S7. Push the new agricultural product wholesale order to the seller's terminal device in real time and display it in real time.

[0068] In step S7, the seller's terminal device may specifically be a computer or tablet in front of the seller, so as to achieve real-time visualization, confirmation and interaction. Furthermore, to further achieve the purpose of intelligent prompt communication, preferably, before pushing the new agricultural product wholesale order to the target salesperson's terminal device in real time, the method also includes, but is not limited to, the following: checking whether the required fields in the new agricultural product wholesale order have been filled in; if not, presenting the required field in the new agricultural product wholesale order in a first warning manner (for example, for key required fields such as "product name" or "quantity," if the dialogue ends and these required fields are still not identified and filled in, reminding the seller to actively inquire with the buyer by highlighting a red border); checking whether the optional fields in the new agricultural product wholesale order have been filled in; if not, presenting the optional fields in the new agricultural product wholesale order in a prompt manner (for example, for optional fields such as "remarks" or "special packaging requirements," if the dialogue ends and these optional fields are still not identified and filled in, gently prompting the seller to actively inquire with the buyer by displaying a blue border); checking whether there are logical errors in the filled fields in the new agricultural product wholesale order; if there are, presenting the filled fields in the new agricultural product wholesale order in a second warning manner. The aforementioned logical errors may include, but are not limited to: mismatched units of quantity ("box" vs. "jin"); prices significantly deviating from the market average; etc. If any of these logical errors are found, immediately mark the corresponding field with a red box and provide a reason.

[0069] Therefore, based on the automatic agricultural product wholesale order generation method described in steps S1 to S7 above, a new automated order generation solution is provided that can separate the voices of both parties in a conversation in real time, deeply understand the semantics of the transaction, and support real-time visual confirmation. First, multiple on-site sound signals are received in real time from the agricultural product sales communication scene by a microphone array. Empirical mode decomposition (EMD), sound source localization, and voiceprint matching technologies are applied in real time to extract the voice signals of the seller and buyer. Then, the extracted results are integrated in real time into a text data stream of the agricultural product sales communication dialogue, and imported into a large language model for semantic understanding and entity information extraction. Next, a matching agricultural product wholesale order template is retrieved based on the extraction results, and the template is filled in and dynamically updated to generate a new agricultural product wholesale order. Finally, the order is pushed to the seller's terminal device in real time for real-time output and display. This improves the efficiency of agricultural product wholesale order generation even in telephone communication scenarios, ensures the accuracy of order information generation, enhances the seller's experience, and facilitates practical application and promotion.

[0070] like Figure 3As shown, the second aspect of this embodiment provides a virtual device for implementing the automatic generation method of agricultural product wholesale orders described in the first aspect, including a sound signal receiving unit, a speech signal extraction unit, a speech-to-text conversion unit, a dialogue text forming unit, a semantic entity extraction unit, a wholesale order generation unit, and an order push display unit that are connected in sequence. The sound signal receiving unit is used to receive multiple on-site sound signals collected in real time by the microphone array at the agricultural product sales communication site, wherein the microphone array includes multiple microphones that correspond one-to-one with the multiple on-site sound signals. The speech signal extraction unit is used to apply empirical mode decomposition technology, sound source localization technology and voiceprint matching technology in real time to extract speech signals from the multiple on-site sound signals to obtain the seller's speech speech signal and the buyer's speech speech signal. The speech-to-text conversion unit is used to process the seller's speech voice signal and the buyer's speech voice signal into speech-to-text in real time to obtain the seller's speech text data stream and the buyer's speech text data stream. The dialogue text forming unit is used to form a real-time dialogue text data stream for agricultural product sales communication based on the seller's speech text data stream and the buyer's speech text data stream. The semantic entity extraction unit is used to import the agricultural product sales communication dialogue text data stream into the first language model in real time for semantic understanding and entity information extraction processing, and obtain semantic understanding and entity information extraction results. The wholesale order generation unit is used to automatically retrieve a matching agricultural product wholesale order template from the order template library based on the semantic understanding and entity information extraction results, and to fill in and dynamically update the content of the agricultural product wholesale order template to generate a new agricultural product wholesale order. The order push and display unit is used to push the new agricultural product wholesale order to the seller's terminal device in real time and display it in real time.

[0071] The working process, working details and technical effects of the aforementioned device provided in the second aspect of this embodiment can be found in the method for automatically generating wholesale orders for agricultural products described in the first aspect, and will not be repeated here.

[0072] like Figure 4As shown, the third aspect of this embodiment provides a computer device for executing the automatic agricultural product wholesale order generation method as described in the first aspect. The device includes a storage module, a processing module, and a transceiver module connected in sequence. The storage module stores a computer program, the transceiver module sends and receives messages, and the processing module reads the computer program and executes the automatic agricultural product wholesale order generation method as described in the first aspect. Specifically, the storage module may include, but is not limited to, random-access memory (RAM), read-only memory (ROM), flash memory, first-in-first-out (FIFO) memory, and / or first-in-last-out (FILO) memory, etc.; the processing module may, but is not limited to, use a microprocessor of the STM32F105 series. Furthermore, the computer device may also include, but is not limited to, a power supply module, a display screen, and other necessary components.

[0073] The working process, working details and technical effects of the aforementioned computer equipment provided in the third aspect of this embodiment can be found in the automatic generation method for agricultural product wholesale orders described in the first aspect, and will not be repeated here.

[0074] This fourth aspect of the embodiment provides a computer-readable storage medium storing instructions comprising the method for automatically generating agricultural product wholesale orders as described in the first aspect. Specifically, the computer-readable storage medium stores instructions that, when executed on a computer, perform the method for automatically generating agricultural product wholesale orders as described in the first aspect. The computer-readable storage medium refers to a data storage medium, and may include, but is not limited to, floppy disks, optical disks, hard disks, flash memory, USB flash drives, and / or Memory Sticks. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0075] The working process, working details and technical effects of the aforementioned computer-readable storage medium provided in the fourth aspect of this embodiment can be found in the method for automatically generating agricultural product wholesale orders as described in the first aspect, and will not be repeated here.

[0076] This fifth aspect of the embodiment provides a computer program product, including a computer program or instructions, which, when executed by a computer, implements the method for automatically generating wholesale agricultural product orders as described in the first aspect. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.

[0077] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for automatically generating wholesale orders for agricultural products, characterized in that, include: Receive multiple on-site sound signals collected in real time from the communication site of agricultural product sales by a microphone array, wherein the microphone array includes multiple microphones that correspond one-to-one with the multiple on-site sound signals; Real-time application of empirical mode decomposition technology, sound source localization technology and voiceprint matching technology is used to extract and process the speech signals of the multiple on-site sound signals to obtain the seller's speech signal and the buyer's speech signal. The seller's speech signal and the buyer's speech signal are processed in real time to convert speech to text, resulting in a seller's speech text data stream and a buyer's speech text data stream. Based on the seller's and buyer's speech text data streams, a real-time agricultural product sales communication dialogue text data stream is generated. The text data stream of the agricultural product sales communication dialogue is imported into the first language model in real time for semantic understanding and entity information extraction processing to obtain the semantic understanding and entity information extraction results. Based on the semantic understanding and entity information extraction results, the matching agricultural product wholesale order template is automatically retrieved from the order template library, and the content of the agricultural product wholesale order template is filled in and dynamically updated to generate a new agricultural product wholesale order. The new agricultural product wholesale orders are pushed to the seller's terminal device in real time and displayed in real time.

2. The method for automatically generating wholesale orders for agricultural products according to claim 1, characterized in that, Real-time application of empirical mode decomposition technology, sound source localization technology, and voiceprint matching technology is used to extract and process the speech signals from the multiple on-site sound signals, resulting in the seller's speech signal and the buyer's speech signal, including: For each of the multiple on-site sound signals, empirical mode decomposition is performed on the corresponding signal in real time to obtain multiple intrinsic mode function components. Based on the multiple intrinsic mode function components, the center frequency points of the corresponding components are determined, wherein the center frequency points of the multiple components correspond one-to-one with the multiple intrinsic mode function components. Based on the multiple component center frequency points of each of the on-site sound signals, at least one adjacent frequency point group is found, wherein the adjacent frequency point group contains at least three component center frequency points corresponding to different speech signals, and the maximum frequency difference of the at least three component center frequency points is less than or equal to a preset frequency threshold. For each adjacent frequency group in the at least one adjacent frequency group, the corresponding sound source location is determined based on the known placement of at least three intrinsic mode function components that correspond one-to-one with the corresponding center frequency points of the at least three components and the known placement of at least three microphones that correspond one-to-one with the at least three intrinsic mode function components. Cluster analysis is performed on the sound source locations of each adjacent frequency group to determine at least two sound source centers; For each of the at least two sound source centers, determine all component center frequencies corresponding to any of the multiple field sound signals from all adjacent frequency point groups whose sound source locations belong to the corresponding center, and reconstruct an independent sound signal from the corresponding center based on all intrinsic mode function components that correspond one-to-one with all component center frequencies. Voiceprint matching is performed on the independent sound signals of each sound source center. If the voiceprint of the independent sound signal of a certain sound source center is found to be the best match with the voiceprint of the target salesperson, then the independent sound signal of that sound source center is taken as the seller's speech voice signal, and the independent sound signal of another sound source center that has a contextual dialogue relationship with the seller's speech voice signal is taken as the buyer's speech voice signal.

3. The method for automatically generating wholesale orders for agricultural products according to claim 2, characterized in that, Based on the center frequency points of the multiple components of the various on-site sound signals, at least one group of adjacent frequency points is found, including: In the frequency domain, examine the center frequency points of the multiple components of each field sound signal in sequence along the direction from small to large frequency. If it is found that the center frequency point of a component of a certain field sound signal is located at the current frequency point and the current frequency point is not within the established frequency domain window, then create a new frequency domain window with the current frequency point as the starting frequency point and the frequency domain width equal to the preset frequency threshold. For each established frequency domain window, determine whether the total number of multiple component center frequency points corresponding to different sound signals within the corresponding window is greater than or equal to 3. If so, include the multiple component center frequency points into an adjacent frequency point group, wherein the adjacent frequency point group contains at least three component center frequency points corresponding to different sound signals, and the maximum frequency difference of the at least three component center frequency points is less than or equal to the preset frequency threshold.

4. The method for automatically generating wholesale orders for agricultural products according to claim 2, characterized in that, For each adjacent frequency group in the at least one adjacent frequency group, the corresponding sound source location is determined based on the known placement of at least three intrinsic mode function components corresponding to the corresponding center frequencies of the at least three components and the known placement of at least three microphones corresponding to the at least three intrinsic mode function components. This includes: For a certain adjacent frequency group in the at least one adjacent frequency group, determine at least three intrinsic mode function components that correspond one-to-one with the corresponding center frequency points of the at least three components; For each pair of intrinsic mode function components in the at least three intrinsic mode function components, the corresponding signal propagation time difference is calculated based on the corresponding two intrinsic mode function components; Based on the known placement positions of at least three microphones that correspond one-to-one with the at least three intrinsic mode function components and the propagation time difference values ​​of each pair of intrinsic mode function components, the sound source location corresponding to a certain adjacent frequency point group is calculated using a time difference localization algorithm.

5. The method for automatically generating wholesale orders for agricultural products according to claim 2, characterized in that, The independent sound signal from another sound source center that has a contextual dialogue relationship with the seller's spoken voice signal is used as the buyer's spoken voice signal, including: The system sequentially traverses all other sound source centers among the at least two sound source centers in order of distance from the aforementioned sound source center: First, it performs speech-to-text processing on the seller's speech voice signal and the independent voice signals of the currently traversed other sound source centers to obtain the seller's speech text data stream and the other speech text data streams. Then, it imports the seller's speech text data stream and the other speech text data streams into the second language model to determine whether there is a contextual dialogue relationship. If there is, the independent voice signal of the currently traversed other sound source center is used as the buyer's speech voice signal; otherwise, it traverses the next other sound source center.

6. The method for automatically generating wholesale orders for agricultural products according to claim 1, characterized in that, The semantic understanding and entity information extraction results include entity information of agricultural product name, entity information of agricultural product specification, entity information of agricultural product quantity, entity information of agricultural product price, entity information of agricultural product transaction time, and / or entity information of agricultural product transaction location. Based on the semantic understanding and entity information extraction results, matching agricultural product wholesale order templates are automatically retrieved from the order template library, including: Based on the agricultural product name entity information in the semantic understanding and entity information extraction results, agricultural product wholesale order templates that match the product category to which the agricultural product name entity information belongs are automatically retrieved from the order template library.

7. The method for automatically generating wholesale orders for agricultural products according to claim 1, characterized in that, Before pushing the new agricultural product wholesale order to the target salesperson's terminal device in real time, the method further includes: Check whether the required fields in the new agricultural product wholesale order have been filled in. If not, display the required field in the new agricultural product wholesale order with a first warning. And / or, check whether the optional fields in the new agricultural product wholesale order have been filled in; if not, present the optional field in the new agricultural product wholesale order in a prompt manner. And / or, check whether there are logical errors in the filled items in the new agricultural product wholesale order; if so, present the filled items in the new agricultural product wholesale order in a second warning manner.

8. An automatic order generation device for agricultural products, characterized in that, It includes a sound signal receiving unit, a speech signal extraction unit, a speech-to-text conversion unit, a dialogue text formation unit, a semantic entity extraction unit, a wholesale order generation unit, and an order push display unit that are connected in sequence. The sound signal receiving unit is used to receive multiple on-site sound signals collected in real time by the microphone array at the agricultural product sales communication site, wherein the microphone array includes multiple microphones that correspond one-to-one with the multiple on-site sound signals. The speech signal extraction unit is used to apply empirical mode decomposition technology, sound source localization technology and voiceprint matching technology in real time to extract speech signals from the multiple on-site sound signals to obtain the seller's speech speech signal and the buyer's speech speech signal. The speech-to-text conversion unit is used to process the seller's speech voice signal and the buyer's speech voice signal into speech-to-text in real time to obtain the seller's speech text data stream and the buyer's speech text data stream. The dialogue text forming unit is used to form a real-time dialogue text data stream for agricultural product sales communication based on the seller's speech text data stream and the buyer's speech text data stream. The semantic entity extraction unit is used to import the agricultural product sales communication dialogue text data stream into the first language model in real time for semantic understanding and entity information extraction processing, and obtain semantic understanding and entity information extraction results. The wholesale order generation unit is used to automatically retrieve a matching agricultural product wholesale order template from the order template library based on the semantic understanding and entity information extraction results, and to fill in and dynamically update the content of the agricultural product wholesale order template to generate a new agricultural product wholesale order. The order push and display unit is used to push the new agricultural product wholesale order to the seller's terminal device in real time and display it in real time.

9. A computer device, characterized in that, The device includes a storage module, a processing module, and a transceiver module that are sequentially connected in communication. The storage module is used to store a computer program, the transceiver module is used to send and receive messages, and the processing module is used to read the computer program and execute the automatic generation method for agricultural product wholesale orders as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that... The computer-readable storage medium stores instructions that, when executed on a computer, perform the method for automatically generating wholesale agricultural product orders as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Agricultural product voice rapid ordering system

    CN106056424A

  • A commodity trading system based on voice interaction

    CN109191269A

  • Monitor trolley

    CN203666729U

  • Business data processing method and device based on artificial intelligence, equipment and medium

    CN117131093A

  • Agricultural product wholesale market voice automatic order generation system based on artificial intelligence

    CN119831707A