Audio processing method, device, electronic device and storage medium
By collecting and identifying audio data during the test ride and determining text data corresponding to the current mode of the vehicle, the problem of sales personnel being difficult to accurately understand user intentions is solved, and a more accurate user intention analysis is achieved.
Patent Information
- Application Number
- CN202210935370.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-04
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-08-04
AI Technical Summary
During the test ride, it is difficult for sales personnel to accurately understand the user's intention to buy a car, resulting in deviations in intention determination.
By collecting audio data between the salesperson and the user, performing voice recognition, text data corresponding to the current mode of the vehicle is determined, and the user's intention is determined based on the text data.
It improves the accuracy of user intentions, reduces the bias in sales personnel's subjective understanding, and allows user intentions to be quantitatively analyzed.
Smart Images

Figure CN115273857B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to the fields of autonomous driving, natural language processing, and speech technology. More specifically, the present disclosure provides an audio processing method, device, electronic device, storage medium, and computer program product. Background Art
[0002] When users want to buy or experience a vehicle, they will go to a car dealer for a test ride and test drive to learn about the vehicle's performance, appearance, price, etc. It can be seen that test rides and test drives are an important link in the user's car purchase decision chain, and are also an important link for car manufacturers to understand user needs, improve their products based on user needs, and carry out digital transformation. Summary of the invention
[0003] The present disclosure provides an audio processing method, an apparatus, an electronic device, a storage medium, and a computer program product.
[0004] According to one aspect of the present disclosure, there is provided an audio processing method, comprising: acquiring audio data, wherein the audio data includes first audio data for a first target object; performing speech recognition on the audio data to obtain text data; determining first text data corresponding to the first audio data from the text data according to a current mode of a vehicle; and determining the intention of the first target object according to the first text data.
[0005] According to another aspect of the present disclosure, an audio processing device is provided, including an acquisition module, a recognition module, a first text determination module, and an intention determination module. The acquisition module is used to acquire audio data, wherein the audio data includes first audio data for a first target object. The recognition module is used to perform speech recognition on the audio data to obtain text data. The first text determination module is used to determine first text data corresponding to the first audio data from the text data according to the current mode of the vehicle. The intention determination module is used to determine the intention of the first target object according to the first text data.
[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method provided by the present disclosure.
[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the method provided by the present disclosure.
[0008] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the method provided by the present disclosure is implemented.
[0009] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0011] Figure 1 is a schematic diagram of an application scenario of the audio processing method and device according to an embodiment of the present disclosure;
[0012] Figure 2 is a schematic flow chart of an audio processing method according to an embodiment of the present disclosure;
[0013] Figure 3 is a schematic flow chart of determining the intention of a first target object according to first text data according to an embodiment of the present disclosure;
[0014] Figure 4 is a schematic flow chart of an audio processing method according to another embodiment of the present disclosure;
[0015] Figure 5 is a schematic diagram of the system architecture of the audio processing method according to an embodiment of the present disclosure;
[0016] Figure 6 is a schematic diagram of an audio processing method according to an embodiment of the present disclosure;
[0017] Figure 7 is a schematic structural block diagram of an audio processing device according to an embodiment of the present disclosure; and
[0018] Figure 8 It is a structural block diagram of an electronic device used to implement the audio processing method of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0019] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0020] During the test ride, users will communicate with sales staff. For example, sales staff will introduce vehicle performance, parameters, price and other information to users, and users will express their evaluation of the vehicle's various performance aspects to sales staff.
[0021] Since the communication between sales staff and users during the test drive is unstructured data, it is difficult to quantify and analyze. Therefore, in some technical solutions, the user's intention is determined manually. For example, the sales staff will understand the user's intention to buy a car based on the communication during the test drive and their personal subjective views.
[0022] It is understandable that sales staff have biased understanding of users’ car-buying intentions based on their subjective understanding, which makes it difficult to accurately determine users’ intentions from communication conversations.
[0023] The disclosed embodiment aims to propose an audio processing method, which collects audio data of a communication conversation between a salesperson and a user, then identifies first text data for the user based on the audio data, and then determines the user's intention based on the first text data. Since the user's intention is analyzed based on the first text, the salesperson's subjective understanding of the user's intention is replaced, so the user's intention can be quantitatively analyzed, thereby improving the accuracy of determining the intention.
[0024] The technical solution provided by the present disclosure will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0025] Figure 1 It is a schematic diagram of an application scenario of the audio processing method and device according to an embodiment of the present disclosure.
[0026] It should be noted that Figure 1 What is shown is merely an example of a system architecture to which the embodiments of the present disclosure can be applied, in order to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.
[0027] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a first terminal device 101 , a second terminal device 102 , a network 103 and a server 104 .
[0028] The first terminal device 101 may be any electronic device having a display screen and supporting web browsing, including but not limited to a smart phone, a tablet computer, a laptop computer, a desktop computer, etc. The user may use the first terminal device 101 to interact with the server 104 through the network 103 to receive or send messages, etc. For example, the user's driver's license may be scanned by the first terminal device 101, and the driver's license image may be uploaded to the server 104. For another example, the user may select the current mode of the vehicle through the first terminal device 101, and the mode may include a test ride mode, a test drive mode, and an end mode, and then send the current mode to the server 104.
[0029] The second terminal device 102 may be an in-vehicle device, such as a smart rearview mirror and an OBD (On Board Diagnostics). The second terminal device 102 may interact with the server 104 through the network 103 to receive or send messages, etc. For example, the second terminal device 102 may collect audio data of a conversation between a first target object (e.g., a user) and a second target object (e.g., a salesperson), and then perform speech recognition on the audio data, and then send the recognized audio text to the server 104. For another example, the second terminal device 102 sends the collected audio data to the server 104, and the server 104 performs speech recognition to obtain the audio text.
[0030] The network 103 is used to provide a medium for a communication link between the first terminal device 101, the second terminal device 102, and the server 104. The network 103 may include various connection types, such as wired and / or wireless communication links, and the like.
[0031] The server 104 may be a server that provides various services, such as a background management server (only as an example) that provides support for websites browsed by users using the first terminal device 101. The background management server may analyze and process the received data such as user requests, and feed back the processing results (such as the intention of the first target object determined based on the audio data, the evaluation value of the second target object, etc.) to the first terminal device 101.
[0032] It should be noted that the audio processing method provided in the embodiment of the present disclosure can generally be executed by the server 104. Accordingly, the audio processing device provided in the embodiment of the present disclosure can generally be set in the server 104. The audio processing method provided in the embodiment of the present disclosure can also be executed by a server or a server cluster that is different from the server 104 and can communicate with the first terminal device 101, the second terminal device 102 and / or the server 104. Correspondingly, the audio processing device provided in the embodiment of the present disclosure can also be set in a server or a server cluster that is different from the server 104 and can communicate with the first terminal device 101, the second terminal device 102 and / or the server 104.
[0033] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0034] Figure 2 is a schematic flow chart of an audio processing method according to an embodiment of the present disclosure.
[0035] like Figure 2 As shown, the audio processing method 200 may include operations S210 to S240.
[0036] In operation S210 , audio data is acquired, where the audio data includes first audio data for a first target object.
[0037] The audio data may represent the audio of the communication between the first target object and the second target object during the test drive. The first target object may be a user who receives the test drive service, such as a consumer who wants to buy a vehicle, a person who wants to experience the vehicle, etc. The second target object may be a salesperson who introduces vehicle information to the first target object.
[0038] The audio data may include first audio data, which may represent audio spoken by a first target object during a test drive, and may also include second audio data, which may represent audio spoken by a second target object during a test drive.
[0039] In actual applications, after informing the first target object that the communication dialogue during the test ride can be collected and obtaining the first target object's consent, an audio collection device can be used to collect audio data. The audio collection device is the first terminal device or the second terminal device mentioned above.
[0040] In operation S220, speech recognition is performed on the audio data to obtain text data.
[0041] For example, ASR (Automatic Speech Recognition) processing may be performed on the audio data to convert the audio data into text data.
[0042] In operation S230, first text data corresponding to the first audio data is determined from the text data according to the current mode of the vehicle.
[0043] For example, the modes of the vehicle may include a test ride mode and a test drive mode.
[0044] For example, the first text data may represent text data converted from audio data of the first target object's speech. The location of the first target object may be determined according to the current mode of the vehicle, and then the sound from the location is used as the first audio data to determine the first text data.
[0045] In operation S240, the intention of the first target object is determined according to the first text data.
[0046] For example, the first target object's comments on the vehicle may be extracted from the first text data, and then the first target object's intention may be determined based on the content of the comments.
[0047] According to the technical solution provided by the embodiment of the present disclosure, by collecting audio data of the communication and dialogue between the first target object and the second target object, the first text data for the first target object is identified based on the audio data, and then the intention of the first target object is determined based on the first text data. Since the intention of the first target object is analyzed based on the first text, the intention of the first target object can be quantitatively analyzed. Compared with the method of subjectively understanding the intention of the first target object through the second target object, the technical solution provided by the present application can improve the accuracy of determining the intention of the first target object.
[0048] According to another embodiment of the present disclosure, the operation of determining the first text data corresponding to the first audio data from the text data according to the current mode of the vehicle may include the following operations: firstly determining the location of the first target object according to the current mode of the vehicle, and determining the data collected at the location as the first audio data. Then, determining the first text data according to the first audio data.
[0049] For example, in the test ride mode, the first target object will sit in a position other than the main driving position, such as the co-pilot position, and accordingly, the second target object will sit in the main driving position to drive the vehicle. Therefore, the data collected from the co-pilot position of the vehicle in the audio data can be determined as the first audio data. Similarly, the data collected from the main driving position of the vehicle in the audio data can be determined as the second audio data.
[0050] For another example, in the test drive mode, the first target object may sit in the main driving position and drive the vehicle, and the second target object may sit in the co-pilot position. Therefore, the data collected from the main driving position of the vehicle in the audio data may be determined as the first audio data. Similarly, the data collected from the co-pilot position of the vehicle in the audio data may be determined as the second audio data.
[0051] In practical applications, taking the use of the second terminal device described above to collect audio data as an example, two sound receiving holes can be set on the second terminal device, with the first sound receiving hole facing the main driving position and the second sound receiving hole facing the co-pilot position. The first sound receiving hole can be used to collect audio from the main driving position, and the second sound receiving hole can be used to collect audio from the co-pilot position.
[0052] For example, the first audio data may be subjected to speech recognition to obtain the first text data. For another example, the first text data may be determined from the text data according to the start time and the end time of the first audio data.
[0053] According to the technical solution provided by the embodiment of the present disclosure, since the position of the first target object is determined according to the current mode of the vehicle and then the first text data is determined according to the position, the first text data corresponding to the first target object is accurately determined, thereby improving the accuracy of the intention.
[0054] Figure 3 is a schematic flowchart of determining the intention of a first target object according to first text data according to an embodiment of the present disclosure.
[0055] like Figure 3 As shown, the method 340 for determining the intention of the first target object according to the first text data may include operations S341 to S344.
[0056] In operation S341, a plurality of target data included in the first text data is determined according to a plurality of predetermined tags, each target data including a target tag for characterizing vehicle performance and emotional text data for characterizing an emotional tendency of the first target object.
[0057] For example, the predetermined tags can be used to evaluate vehicle performance, and multiple predetermined tags can be pre-set according to needs. For example, "maintenance", "fuel consumption", "oil brand" and the like can be used as predetermined tags for evaluating vehicle economy, and "front row space", "vehicle width", "wheelbase" and the like can be used as predetermined tags for evaluating vehicle space information.
[0058] The target tag may refer to a tag that is included in the first text data and is the same as or similar to the predetermined tag.
[0059] Sentimental tendencies may include positive, neutral, and negative. Sentimental text data is text used to describe sentimental tendencies. For example, sentimental text data such as "comfortable", "cheap", and "spacious" can describe positive sentiments. Sentimental text data such as "normal level", "ordinary", "almost", and "okay" can describe neutral sentiments. Sentimental text data such as "uncomfortable", "too expensive", and "crowded" can describe negative sentiments.
[0060] In one example, a target tag matching multiple predetermined tags can be first determined from the first text data, and the target tag includes a noun. For example, the first text data can be segmented to obtain multiple first participles, and then the similarity between the first participle and the predetermined tag is calculated, and the first participle whose similarity is greater than or equal to the first threshold is determined as the target tag. Then, the context text data corresponding to the target tag can be determined from the first text data, for example, the sentence where the target tag is located is determined as the context text data. Then, the target adjective in the context text data can be determined, for example, the part of speech of the first participle contained in the context text data can be identified, and the target adjective can be determined therefrom. Next, the target tag and the target adjective can be determined as the target data. Using the technical solution provided by this example, the emotional tendency of the first target object towards the predetermined vehicle performance can be accurately determined through the name and adjective, thereby improving the accuracy of the schematic diagram.
[0061] In another example, an NLP (natural language processing) model trained in advance with automotive industry-related corpus can be used to process the first text data, thereby obtaining the target labels contained in each sentence spoken by the first target object and the sentiment tendency of each label.
[0062] In operation S342, target data including positive sentiment text data among the plurality of target data is determined as first target data.
[0063] For example, the first target data may include "cheap maintenance", "spacious front-seat space", "low fuel consumption", etc.
[0064] In operation S343, target data including negative sentiment text data among the plurality of target data is determined as second target data.
[0065] For example, the second target data may include "expensive maintenance", "crowded front space", "high fuel consumption", etc.
[0066] In operation S344, an intention is determined based on the first target data and the second target data.
[0067] For example, if the number of the first target data is greater than the number of the second target data, it is determined that the first target object is relatively satisfied with the test-driving vehicle and has a greater intention to purchase the vehicle. If the number of the first target data is less than or equal to the number of the second target data, it is determined that the first target object has a small intention to purchase the vehicle or has no intention to purchase the vehicle.
[0068] For another example, when the number of the first target data is greater than or equal to the first predetermined number, the number of the second target data is less than or equal to the second predetermined number, and the first predetermined number is greater than the second predetermined number, it can be determined that the first target object has a greater intention to purchase a car. The first predetermined number can be 15, and the second predetermined number can be 5.
[0069] According to the technical solution provided by the embodiments of the present disclosure, since the first target data and the second target data are determined based on the positive emotional text data and the negative emotional text data, and the intention is determined based on the first target data and the second target data, the intention of the first target object can be quantitatively analyzed, thereby improving the accuracy of the intention.
[0070] According to another embodiment of the present disclosure, multiple predetermined tags can be divided into multiple tag sets in advance, each tag set includes at least two tags with a hierarchical relationship, and the method may also include the following operations: after determining multiple target data included in the first text data, the tag set where the target tag is located is determined as the target tag set, and then the preference information of the first target object is determined based on the target tag set and the hierarchical relationship.
[0071] For example, in the first tag set, "economy" is a first-level tag, the next-level tags of the "economy" tag include "budget" and "fuel consumption", the next-level tags of the "budget" tag include "maintenance costs" and "engine oil brand", and the next-level tags of the "fuel consumption" tag include "fuel-saving mode", "fuel consumption per 100 kilometers", "instantaneous fuel consumption" and "constant speed fuel consumption".
[0072] For example, in the second tag set, "space" is a first-level tag, the next-level tags of the "space" tag include "front row space", "second row space" and "physical size", the next-level tags of the "front row space" tag include "main driver space" and "driving seat space", the next-level tags of the "second row space" tag include "rear top leg", "middle bulge" and "second row middle seat space", and the next-level tags of the "physical size" tag include "vehicle width", "width" and "wheelbase".
[0073] For example, when it is determined that the target tag in the first text data includes "fuel consumption per 100 kilometers", the above-mentioned first tag set can be determined as the target tag set. Then, other tags that have a hierarchical relationship with the target tag can be determined from the target tag set, such as determining the "fuel consumption" tag and the "economy" tag from the above-mentioned first tag set. Through these tags, it can be determined that the first target object is concerned about the economy of the vehicle, and is more concerned about fuel consumption in economy, especially fuel consumption per 100 kilometers. Based on this focus, it can be determined that the first target object likes vehicles with high economy, and then a car model that meets his or her preferences can be recommended to the first target object.
[0074] The disclosed embodiment determines the preference information of the first target object according to the hierarchical relationship and the tag set, and then recommends the preferred car model to the first target object, thereby improving the test ride and test drive experience.
[0075] Figure 4 is a schematic flowchart of an audio processing method according to another embodiment of the present disclosure.
[0076] like Figure 4 As shown, the audio processing method 400 may further include operations S450 to S470.
[0077] In operation S450, second text data corresponding to the second audio data is determined from the text data.
[0078] For example, the location of the second target object may be determined first according to the current mode of the vehicle, and the data collected at the location may be determined as the second audio data.
[0079] In operation S460, at least one target corpus matching the second text data is determined according to a corpus set including at least one predetermined corpus.
[0080] For example, multiple sales scripts can be pre-configured based on vehicle models. Each sales script can include script content and predetermined corpus. The script content can be a text introducing vehicle performance, and the predetermined corpus can be keywords in the script content.
[0081] For example, the second text data may be segmented to obtain multiple second segmented words, and then the similarity between the second segmented words and the predetermined corpus is calculated, where the similarity may be text similarity, and then the predetermined corpus with a similarity greater than a second threshold is determined as the target corpus.
[0082] In operation S470 , an evaluation value for a second target object is determined according to the at least one target corpus.
[0083] According to the technical solution provided in the embodiment of the present disclosure, the situation of the second target object introducing and explaining the vehicle information can be evaluated to determine whether the second target object explains the predetermined corpus to the first target object, such as whether to introduce specific performance parameters and functions of the vehicle to the first target object. In this way, the working condition of the second target object can be evaluated, the working behavior of the second target object can be constrained, and the test ride effect can be ensured.
[0084] In addition, in some embodiments, a prompt may be given for the predetermined corpus in the corpus set that does not match the second text data.
[0085] According to another embodiment of the present disclosure, the corpus set includes a plurality of corpus subsets, each corpus subset includes at least one predetermined corpus related to the same vehicle performance.
[0086] For example, a certain speech content related to vehicle power is "the vehicle is equipped with a 2.0L naturally aspirated engine with a maximum power of 126kW and a maximum torque of 209N·m". The corpus subset corresponding to the speech content may include predetermined corpora such as "2.0L engine" and "strong power".
[0087] For another example, a certain speech content related to vehicle space is "There is a lot of space inside the car, with a wheelbase of 2.75 meters, the largest in the same class, and up to 980mm of legroom, giving people more room to move around." The corpus subset corresponding to the speech content may include predetermined corpora such as "large space", "wheelbase 2750mm", "wheelbase 2.75m", and "legroom 980mm".
[0088] Accordingly, the operation of determining the evaluation value for the second target object may include the following operations: determining at least one target corpus subset in which the at least one target corpus is located according to the at least one target corpus, and then determining the evaluation value according to the predetermined evaluation value corresponding to each of the at least one target corpus subsets.
[0089] For example, the corpus set includes corpus subsets a, b, and c, wherein the predetermined corpus in corpus subset a includes "2.0L engine" and "strong power", the predetermined corpus in corpus subset b includes "large space", "wheelbase 2750mm", "wheelbase 2.75m", and "leg space 980mm", and the predetermined corpus in corpus subset c includes "low fuel consumption" and "cheap maintenance".
[0090] When the target corpus includes "strong power", "large space" and "wheelbase 2.75m", the target corpus subset includes corpus subsets a and b.
[0091] For example, the predetermined evaluation value corresponding to the corpus subset a is added to the predetermined evaluation value corresponding to the corpus subset b, and the sum of the two is used as the evaluation value of the second target object.
[0092] It can be seen that the second text data matches any predetermined corpus in the corpus subset, indicating that the second target object introduces the vehicle performance corresponding to the corpus subset.
[0093] By adopting the technical solution provided by the embodiment of the present disclosure, it is possible to accurately determine which vehicle performance the second target object explains. At the same time, the second target object does not need to introduce multiple predetermined corpora contained in the corpus subset related to the same vehicle performance. Therefore, the workload of the second target object can be reduced while ensuring the service quality.
[0094] Figure 5 is a schematic diagram of the system architecture of the audio processing method according to an embodiment of the present disclosure.
[0095] like Figure 5 As shown, this embodiment involves a vehicle terminal 510 , a client terminal 520 and a cloud terminal 530 .
[0096] The vehicle terminal 510 may adopt the second terminal device described above, for example, the vehicle terminal 510 may include an OBD and a smart rearview mirror.
[0097] OBD can obtain vehicle information such as VIN code (vehicle frame number), single mileage, speed, etc., and then send the collected vehicle information to the smart rearview mirror.
[0098] The smart rearview mirror is installed inside the vehicle, and is used as a rearview mirror on the one hand, and can transmit data on the other hand. For example, the smart rearview mirror can collect audio data, and can upload the collected audio data and the vehicle information received from the OBD to the cloud 530, for example, 10 seconds of audio data is uploaded as a slice to the cloud 530, and GPS data, vehicle speed and other data are uploaded every 2 seconds. For another example, the smart rearview mirror can bind its own identification code (such as SN code), OBD identification code and VIN code to three identification codes, and send the three identification codes to the cloud 530. For another example, the smart rearview mirror can voice broadcast the current vehicle mode during the test ride. For another example, the smart rearview mirror can recommend and explain the functions of the vehicle according to predetermined rules. For example, when the vehicle is detected to have traveled for a predetermined time and to a predetermined position, it can voice broadcast relevant information such as the vehicle's comfortable suspension, reasonable turning radius, and high fuel economy, so that the first target object can understand the vehicle information and improve the test ride experience.
[0099] The client 520 can adopt the first terminal device described above, and the client 520 can install an application, which can provide relevant test ride and test drive functions such as scanning code to start the vehicle, information collection, electronic agreement signing, vehicle mode switching, etc., thereby simplifying the test ride and test drive operation process.
[0100] For example, the first process can be started by scanning the test drive identification code using the client 520. For another example, after the first process is started, the information collection page (such as a scanning page) is displayed through the client 520, and the driver's license image of the first target object is collected through the information collection page, and then the client 520 can perform OCR (Optical Character Recognition) recognition on the driver's license image, or the client 520 uploads the driver's license image to the cloud 530 and the cloud 530 performs OCR recognition. For another example, the electronic agreement for the test drive is displayed through the client 520, and the second process is triggered after receiving the electronic signature of the first target object. For another example, a mode button can be displayed after the second process is started, and the vehicle modes such as test ride, test drive, and end can be selected through the mode button.
[0101] The cloud 530 can use the server described above. The cloud 530 can perform relevant data transmission, data processing, status detection and command transmission with the client 520 and the vehicle terminal 510, thereby realizing basic data management, basic service management, business service management and other functions.
[0102] For example, the cloud 530 can manage basic data such as vehicles, models, dealers, second target objects, test drive routes, etc. The basic service management function can be implemented based on MQTT (message queue telemetry transmission), TSDB (time series database), OCR (optical character recognition), ASR (automatic speech recognition), NLP (natural language processing), BOS (object storage), etc. For example, the cloud 530 can implement real-time message reminders during the test drive based on MQTT long connection, such as sending message events during the test drive to the car terminal 510 in real time, so that the car terminal 510 can perform voice broadcast. For another example, GPS data storage and millisecond-level query can be implemented based on TSDB. The driver's license image can be recognized based on OCR, and the audio data can be converted into audio text based on ASR. For another example, the cloud 530 can use BOS to store audio data during the test drive. For another example, the cloud 530 can bind SN code, VIN and OBD identification code. For another example, a test drive related data analysis report can also be generated based on existing data. The business service may include determining the intention of the first target object and the evaluation value of the second target object by using the audio text, and the business service may be implemented based on NLP in the basic service function.
[0103] The technical solution provided by the disclosed embodiment can alleviate the problems of cumbersome test ride and test drive processes, the inability to standardize the sales language of the second target object, and the inability to timely collect and accurately analyze customers' car purchase intentions.
[0104] Figure 6It is a schematic diagram of an audio processing method according to an embodiment of the present disclosure.
[0105] like Figure 6 As shown, this embodiment involves a vehicle terminal 610 , a client terminal 620 and a cloud terminal 630 .
[0106] The vehicle terminal 610 may include an OBD and a smart rearview mirror. The smart rearview mirror reads the vehicle VIN code through the OBD, and then the smart rearview mirror sends the smart rearview mirror identification code (such as SN code), the vehicle VIN code and the OBD identification code to the cloud 630. In addition, a long connection can be established between the smart rearview mirror and the cloud 630.
[0107] The cloud 630 determines whether the identification code (e.g., SN code) of the smart rearview mirror, the vehicle VIN code, and the OBD identification code are bound. For example, the cloud 630 determines whether the three identification codes are bound by detecting whether an association has been established between the three identification codes in the database. If bound, the cloud 630 outputs the test drive identification code corresponding to the three identification codes. If not bound, the cloud 630 binds the three identification codes, updates the data in the database, and then outputs the test drive identification code for display by the vehicle terminal 620.
[0108] After the first target object and the second target object get on the vehicle, the test drive identification code can be scanned through the client 620, and then the page will be jumped to the information collection page, and the driver's license of the first target object can be scanned through the page to obtain the driver's license image, and then the client 620 performs OCR recognition on the driver's license image, or sends the driver's license image to the cloud 630, which performs OCR recognition to obtain the information of the first target object. Since there is no need to manually fill in information in the paper application form, the test drive process can be simplified.
[0109] The client 620 can display the test drive agreement, and after receiving the electronic signature of the first target object, the first target object or the second target object can select the vehicle mode through the mode button displayed by the client 620, thereby facilitating the first target object or the second target object to select the current mode.
[0110] After the client 620 detects that the mode button is triggered, the client 620 can call the cloud 630 interface, the cloud 630 determines the current mode of the vehicle, and controls the vehicle end 610 to voice broadcast the current mode of the vehicle through MQTT asynchronous messages. For example, if the client 620 detects that the test drive mode button is triggered, the cloud 630 will control the vehicle end to voice broadcast the test drive mode. For another example, after the client 620 detects that the current mode is switched, the client 620 sends a message indicating the current mode to the cloud 630, and the cloud 630 controls the vehicle end 610 to voice broadcast the switched mode based on the message through MQTT asynchronous messages. For another example, after the client 620 detects that the end button is triggered, the cloud 630 updates the current mode and then controls the vehicle end 610 to voice broadcast the end of the test drive.
[0111] During the test drive, the smart rearview mirror of the vehicle terminal 610 can collect audio data of the communication between the first target object and the second target object, and then perform ASR recognition on the audio data to obtain audio text. Alternatively, the vehicle terminal 610 sends the audio data to the cloud 630, and the cloud 630 performs ASR recognition on the audio data to obtain audio text. Then the vehicle terminal 610 sends the audio data and audio text to the cloud 630.
[0112] The cloud 630 may process the audio data, for example, by merging multiple audio clips into one audio file, and then storing the audio file based on the BOS.
[0113] The cloud 630 can extract the first text data of the first target object and the second text data of the second target object from the audio text. For the first text data, the emotion of the first target object can be analyzed based on the predetermined tag and the context of the predetermined tag, and then the intention of the first target object can be determined based on the emotion of the first target object, for example, whether the first target object intends to buy a car. For the second text data, the cloud 630 can determine the evaluation value of the second target object based on the predetermined corpus. The cloud 630 can also display an analysis report, which may include the intention of the first target object, the evaluation value of the second target object, the predetermined corpus not mentioned by the second target object, etc. The cloud 630 can also convert the audio text into a dialogue form on the cloud 630, and then display it through the client 610.
[0114] The cloud 630 can also automatically end the test drive. For example, if the cloud 630 detects that the vehicle has not moved within a predetermined time period, the current mode of the vehicle can be updated to an end mode. The predetermined time period can be half an hour.
[0115] The cloud 630 can also generate a record of the test drive, which may include: the distance, start time, end time, dealership information, etc. of the test drive.
[0116] Figure 7 is a schematic structural block diagram of an audio processing device according to an embodiment of the present disclosure.
[0117] like Figure 7 As shown, the audio processing device 700 may include an acquisition module 710 , a recognition module 720 , a first text determination module 730 , and an intention determination module 740 .
[0118] The acquisition module 710 is used to acquire audio data, wherein the audio data includes first audio data for a first target object.
[0119] The recognition module 720 is used to perform speech recognition on the audio data to obtain text data.
[0120] The first text determination module 730 is used to determine first text data corresponding to the first audio data from the text data according to the current mode of the vehicle.
[0121] The intention determination module 740 is used to determine the intention of the first target object according to the first text data.
[0122] According to another embodiment of the present disclosure, the first text determination module includes a first determination submodule, a second determination submodule and a third determination submodule. The first determination submodule is used to determine the data collected for the co-pilot position of the vehicle in the audio data as the first audio data when the current mode is the test drive mode. The second determination submodule is used to determine the data collected for the main driver position of the vehicle in the audio data as the first audio data when the current mode is the test drive mode. The third determination submodule is used to determine the first text data based on the first audio data.
[0123] According to another embodiment of the present disclosure, the intention determination module includes a fourth determination submodule, a fifth determination submodule, a sixth determination submodule and a seventh determination submodule. The fourth determination submodule is used to determine multiple target data included in the first text data based on multiple predetermined labels, each target data including a target label for characterizing vehicle performance and emotional text data for characterizing the emotional tendency of the first target object. The fifth determination submodule is used to determine the target data including positive emotional text data among the multiple target data as the first target data. The sixth determination submodule is used to determine the target data including negative emotional text data among the multiple target data as the second target data. The seventh determination submodule is used to determine the intention based on the first target data and the second target data.
[0124] According to another embodiment of the present disclosure, a plurality of predetermined tags are divided into a plurality of tag sets, each tag set including at least two tags having a hierarchical relationship. The above-mentioned audio processing device also includes a tag set determination module and a preference determination module. The tag set determination module is used to determine the tag set where the target tag is located as the target tag set after determining the plurality of target data included in the first text data. The preference determination module is used to determine the preference information of the first target object based on the target tag set and the hierarchical relationship.
[0125] According to another embodiment of the present disclosure, the fourth determination submodule includes a first determination unit, a second determination unit, a third determination unit, and a fourth determination unit. The first determination unit is used to determine a target tag matching a plurality of predetermined tags from the first text data, wherein the target tag includes a noun. The second determination unit is used to determine context text data corresponding to the target tag from the first text data. The third determination unit is used to determine a target adjective in the context text data. The fourth determination unit is used to determine the target tag and the target adjective as target data.
[0126] According to another embodiment of the present disclosure, the audio data also includes second audio data for a second target object. The above-mentioned audio processing device also includes a second text determination module, a corpus determination module and an evaluation value determination module. The second text determination module is used to determine the second text data corresponding to the second audio data from the text data. The corpus determination module is used to determine at least one target corpus matching the second text data based on a corpus set including at least one predetermined corpus. The evaluation value determination module is used to determine the evaluation value for the second target object based on at least one target corpus.
[0127] According to another embodiment of the present disclosure, the corpus set includes a plurality of corpus subsets, each corpus subset includes at least one predetermined corpus related to the same vehicle performance, and each corpus subset corresponds to a predetermined evaluation value. The evaluation value determination module includes a subset determination submodule and an evaluation value determination submodule. The subset determination submodule is used to determine at least one target corpus subset in which at least one target corpus is located based on at least one target corpus. The evaluation value determination submodule is used to determine the evaluation value based on the predetermined evaluation value corresponding to each of the at least one target corpus subsets.
[0128] In the technical solution of this disclosure, the collection, storage, use, processing, transmission, provision and disclosure of the target object's personal information are in compliance with the relevant laws and regulations and do not violate public order and good morals. The target object can be the first target object and the second target object mentioned above.
[0129] In the technical solution of the present disclosure, the authorization or consent of the target object is obtained before obtaining or collecting the personal information of the target object.
[0130] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, comprising at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned audio processing method.
[0131] According to an embodiment of the present disclosure, the present disclosure further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the above audio processing method.
[0132] According to an embodiment of the present disclosure, the present disclosure further provides a computer program product, including a computer program, and the computer program implements the above audio processing method when executed by a processor.
[0133] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0134] like Figure 8 As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0135] A number of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0136] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as an audio processing method. For example, in some embodiments, the audio processing method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the audio processing method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the audio processing method in any other appropriate manner (e.g., by means of firmware).
[0137] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0138] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0139] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0140] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0141] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0142] A computer system may include clients and servers. Clients and servers are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship to each other.
[0143] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0144] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. An audio processing method, comprising: Acquire audio data, wherein the audio data includes first audio data for a first target object; Performing speech recognition on the audio data to obtain text data; Determining first text data corresponding to the first audio data from the text data according to a current mode of the vehicle, including: determining a location of the first target object according to the current mode, determining data collected at the location of the first target object as the first audio data, and determining the first text data according to the first audio data; and The intention of the first target object is determined according to the first text data.
2. The method according to claim 1, wherein: The step of determining the position of the first target object according to the current mode and determining the data collected for the position of the first target object as the first audio data includes: When the current mode is a test ride mode, determining data collected from the co-pilot position of the vehicle in the audio data as the first audio data; and When the current mode is the test drive mode, data collected from the main driving position of the vehicle in the audio data is determined as the first audio data.
3. The method according to claim 1, wherein: Determining the intention of the first target object according to the first text data includes: According to a plurality of predetermined tags, determining a plurality of target data included in the first text data, each target data including a target tag for characterizing vehicle performance and emotional text data for characterizing the emotional tendency of the first target object; Determine the target data including the positive sentiment text data among the plurality of target data as the first target data; determining, among the plurality of target data, target data including negative sentiment text data as second target data; and The intention is determined based on the first target data and the second target data.
4. The method according to claim 3, wherein: The plurality of predetermined tags are divided into a plurality of tag sets, each tag set including at least two tags having a hierarchical relationship; the method further includes: after determining a plurality of target data included in the first text data, Determine the tag set where the target tag is located as the target tag set; and Determine the preference information of the first target object according to the target tag set and the hierarchical relationship.
5. The method according to claim 3, wherein: The step of determining, according to the plurality of predetermined tags, the plurality of target data included in the first text data comprises: Determining, from the first text data, a target tag that matches the plurality of predetermined tags, the target tag comprising a noun; Determining context text data corresponding to the target tag from the first text data; determining a target adjective in the contextual text data; and The target label and the target adjective are determined as the target data.
6. The method according to any one of claims 1 to 5, wherein: The audio data further includes second audio data for a second target object; and the method further includes: Determining second text data corresponding to the second audio data from the text data; Determining at least one target corpus matching the second text data according to a corpus set including at least one predetermined corpus; and An evaluation value for the second target object is determined according to the at least one target corpus.
7. The method according to claim 6, wherein: The corpus set includes a plurality of corpus subsets, each corpus subset includes at least one predetermined corpus related to the same vehicle performance, and each corpus subset corresponds to a predetermined evaluation value; The determining, according to the at least one target corpus, an evaluation value for the second target object comprises: Determining, according to the at least one target corpus, at least one target corpus subset to which the at least one target corpus belongs; as well as The evaluation value is determined according to a predetermined evaluation value corresponding to each of the at least one target corpus subsets.
8. An audio processing device, comprising: An acquisition module, configured to acquire audio data, wherein the audio data includes first audio data for a first target object; A recognition module, used for performing speech recognition on the audio data to obtain text data; a first text determination module, configured to determine first text data corresponding to the first audio data from the text data according to a current mode of the vehicle, comprising: determining a location of the first target object according to the current mode, determining data collected at the location of the first target object as the first audio data, and determining the first text data according to the first audio data; and An intention determination module is used to determine the intention of the first target object based on the first text data.
9. The device according to claim 8, wherein: The first text determination module comprises: A first determining submodule is configured to determine, when the current mode is a test ride mode, data collected from a passenger seat of the vehicle in the audio data as the first audio data; and The second determining submodule is used to determine, when the current mode is the test drive mode, data collected from the main driving position of the vehicle in the audio data as the first audio data.
10. The device according to claim 8, wherein: The intention determination module comprises: A fourth determination submodule, configured to determine, based on a plurality of predetermined labels, a plurality of target data included in the first text data, each target data including a target label for characterizing vehicle performance and emotional text data for characterizing an emotional tendency of the first target object; A fifth determination submodule, configured to determine the target data including the positive sentiment text data among the plurality of target data as the first target data; a sixth determination submodule, configured to determine the target data including the negative sentiment text data among the plurality of target data as the second target data; and A seventh determination submodule is used to determine the intention according to the first target data and the second target data.
11. The device according to claim 10, wherein: The plurality of predetermined tags are divided into a plurality of tag sets, each tag set including at least two tags having a hierarchical relationship; the apparatus further includes: a tag set determining module, configured to determine, after determining a plurality of target data included in the first text data, a tag set in which the target tag is located as a target tag set; and The preference determination module is used to determine the preference information of the first target object according to the target tag set and the hierarchical relationship.
12. The device according to claim 10, wherein: The fourth determining submodule includes: A first determining unit, configured to determine, from the first text data, a target tag matching the plurality of predetermined tags, the target tag comprising a noun; A second determining unit, configured to determine context text data corresponding to the target tag from the first text data; A third determining unit, configured to determine a target adjective in the context text data; and The fourth determining unit is used to determine the target tag and the target adjective as the target data.
13. The device according to any one of claims 8 to 12, wherein: The audio data further includes second audio data for a second target object; and the device further includes: A second text determination module, used to determine second text data corresponding to the second audio data from the text data; a corpus determination module, configured to determine at least one target corpus matching the second text data according to a corpus set including at least one predetermined corpus; and An evaluation value determination module is used to determine an evaluation value for the second target object based on the at least one target corpus.
14. The device according to claim 13, wherein: The corpus set includes a plurality of corpus subsets, each corpus subset includes at least one predetermined corpus related to the same vehicle performance, and each corpus subset corresponds to a predetermined evaluation value; The evaluation value determination module comprises: A subset determination submodule, configured to determine, based on the at least one target corpus, at least one target corpus subset to which the at least one target corpus belongs; as well as The evaluation value determination submodule is used to determine the evaluation value according to the predetermined evaluation value corresponding to each of the at least one target corpus subsets.
15. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.
17. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Audio processing method and device based on automatic speech recognition
CN115482823A