A method and apparatus for screening fraudulent order processing based on voiceprint and ASR technology

By using voiceprint and ASR technologies to conduct multi-dimensional analysis of ride-hailing order recording data, the problem of low efficiency and accuracy in identifying ride-hailing drivers' empty-running and fraudulent order behavior has been solved. This has enabled efficient and accurate identification of fraudulent order risks, reduced the false judgment rate and platform losses, and improved the driver experience.

CN113919909BActive Publication Date: 2025-12-02广州宸祺出行科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111183774.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-11
Publication Date
2025-12-02
Estimated Expiration
2041-10-11

AI Technical Summary

Technical Problem

In existing technologies, it is difficult to efficiently and accurately identify the behavior of ride-hailing drivers who drive empty and fraudulently obtain orders, resulting in large errors in judging the operational status of ride-hailing platforms, affecting driver experience and causing losses to the platform.

Method used

Voiceprint and ASR technologies are used to analyze the trip recording data of ride-hailing orders. By using a series of models for identifying empty vehicle scenarios, turn signal scenarios, in-vehicle passenger scenarios, and navigation sound scenarios, scene tags are generated and combined with a fraudulent order identification model to achieve multi-dimensional fraudulent order risk analysis.

Benefits of technology

It improved the efficiency and accuracy of identifying fraudulent ride-hailing orders, reduced the false judgment rate, minimized the losses to the platform caused by fraudulent orders, and enhanced the driver experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113919909B_ABST
    Figure CN113919909B_ABST
Patent Text Reader

Abstract

This invention discloses a method and apparatus for screening ride-hailing orders for fraudulent transactions based on voiceprint and ASR (Automatic Recognition) technology. The method includes: acquiring trip recording data corresponding to ride-hailing orders; performing voiceprint analysis on the trip recording data to obtain voiceprint data; processing the trip recording data based on ASR technology to obtain speech-text data; inputting the trip recording data, voiceprint data, and speech-text data into a series of interconnected empty-vehicle scene recognition models, turn signal scene recognition models, in-vehicle passenger scene recognition models, and navigation sound scene recognition models to obtain scene recognition results; generating corresponding scene tags for ride-hailing orders with fraudulent transaction risks based on the scene recognition results; inputting the scene tags and ride-hailing order data into the fraudulent transaction identification model for fraudulent transaction identification, and outputting corresponding fraudulent transaction risk tags based on the fraudulent transaction identification results. This invention achieves efficient and accurate fraudulent transaction identification and screening through multi-dimensional recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of screening methods for ride-hailing fraud, specifically to a method and apparatus for screening fraudulent ride-hailing services based on voiceprint and ASR technology. Background Technology

[0002] With the development of the internet, people's travel methods have also changed. Among them, ride-hailing services have been continuously promoted due to their convenience and timeliness. In the process of promoting ride-hailing services, ride-hailing platforms usually provide subsidies to drivers during peak hours or when there are daily order limits, or subsidies are issued based on the number of orders and performance. Typically, the amount of subsidy issued to drivers is determined based on the ride-hailing order situation.

[0003] However, some ride-hailing drivers resort to various cheating methods to obtain subsidies, the most common being empty runs to generate fake orders, leading to abnormal operation of ride-hailing services. In traditional online ride-hailing order-boosting screening and investigation, determining whether a driver has fraudulently generated orders through empty runs requires analyzing completed orders. Monitoring drivers is difficult, and existing order analysis methods often contain errors, with inadequate evidence collection methods making it impossible to obtain direct proof, significantly impacting the assessment of ride-hailing service operation status.

[0004] After researching the ride-hailing industry, the applicant discovered that existing methods for identifying ride-hailing drivers engaging in fraudulent order placement are labor-intensive, resulting in low screening efficiency and a high risk of errors, negatively impacting driver experience. Therefore, there is an urgent need to invent a highly efficient and accurate method for screening ride-hailing drivers who are engaging in fraudulent order placement. Summary of the Invention

[0005] In order to overcome the technical defects of existing ride-hailing order-brushing identification methods, which are characterized by low efficiency and low accuracy, this invention provides a method and apparatus for screening empty ride-hailing order-brushing based on voiceprint and ASR technology.

[0006] To solve the above problems, the present invention is implemented according to the following technical solution:

[0007] In a first aspect, this invention discloses a method for screening fraudulent order placement based on voiceprint and ASR technology, comprising the following steps:

[0008] Obtain the trip recording data corresponding to the ride-hailing order;

[0009] Perform voiceprint analysis on the trip recording data to obtain voiceprint data;

[0010] The trip recording data is processed using ASR technology to obtain voice-text data;

[0011] The trip recording data, voiceprint data, and voice text data are respectively input into the cascaded empty scene recognition model, turn signal scene recognition model, in-vehicle occupant scene recognition model, and navigation sound scene recognition model to obtain scene recognition results.

[0012] Based on the scene recognition results, corresponding scene tags are generated for ride-hailing orders that have the risk of fraudulent orders;

[0013] The scene tags and ride-hailing order data corresponding to the ride-hailing order are input into the order fraud identification model for order fraud identification. Based on the order fraud identification results, the corresponding order fraud risk tags are output.

[0014] As a preferred embodiment, obtaining the scene recognition result specifically includes:

[0015] The scene recognition model is activated by inputting trip recording data, voiceprint data, and speech-text data into multiple scene recognition models running in series. These models include an empty vehicle scene recognition model, a turn signal scene recognition model, an in-vehicle occupant scene recognition model, and a navigation sound scene recognition model. The empty vehicle scene recognition model includes several models for identifying various empty vehicle conditions, performing semantic analysis based on trip recording and speech-text data to determine the empty vehicle status. The turn signal scene recognition model includes several turn signal models for identifying different turn signal conditions, performing semantic analysis based on the voiceprint and voiceprint data of the turn signal prompt sounds. The system performs a line matching comparison to obtain the turn signal status; the in-vehicle occupancy scene recognition model includes several passenger number models, analyzes the number of human voiceprints in the trip recording data and voiceprint data, and performs semantic analysis on the speech text data to obtain the number of people in the vehicle; the navigation sound scene recognition model includes several navigation sound models, performs voiceprint analysis on the voiceprint data based on the navigation sound, and performs semantic analysis on the corresponding speech text data of the navigation sound to obtain the navigation status; after parallel processing of empty vehicle status, turn signal status, passenger number status, and navigation status, the corresponding scene recognition results are output.

[0016] As a preferred embodiment, the step of inputting the scene tag corresponding to the ride-hailing order and the ride-hailing order data into the order fraud identification model for order fraud identification, and outputting the corresponding order fraud risk tag based on the order fraud identification result, specifically includes:

[0017] After obtaining the scene tags generated from the scene recognition results of multiple serially running scene recognition models, the fraudulent order recognition model is run. The scene tags, as well as the voiceprint data and voice-text data corresponding to the ride-hailing order, are input into the fraudulent order recognition model. The fraudulent order recognition model performs fraudulent order recognition based on the voiceprint data and voice-text data. Combined with the ride-hailing order data, the scene tags are re-verified and corrected, and the corresponding fraudulent order risk tags are generated and output. The fraudulent order risk tags are then added to the corresponding ride-hailing order.

[0018] As a preferred embodiment, the step of performing voiceprint analysis on the trip recording data to obtain voiceprint data specifically includes:

[0019] The voiceprints of drivers corresponding to ride-hailing orders are pre-registered in the voiceprint database. Trip recording data is input into the voiceprint model to analyze the number and type of voiceprints in the trip recordings. The voiceprints of turn signal prompts, navigation sounds, and human voices are extracted. The driver's voiceprint is removed from the human voices through voiceprint matching to obtain the passenger's voiceprint. The voiceprints of turn signal prompts, navigation sounds, and passenger voices are then integrated into voiceprint data.

[0020] As a preferred embodiment, the process of processing the trip recording data based on ASR technology to obtain voice-text data specifically includes:

[0021] The trip recording data is input into the ASR model, which converts the trip recording data into corresponding text to form speech-text data.

[0022] Secondly, this invention also discloses a device for screening fraudulent order-taking based on voiceprint and ASR technology, comprising a recording acquisition module, a voiceprint analysis module, a text conversion module, a scene recognition module, a scene tagging module, and an order-taking identification module, specifically including:

[0023] The audio recording acquisition module is used to acquire trip audio recording data corresponding to ride-hailing orders;

[0024] The voiceprint analysis module is used to perform voiceprint analysis on the travel recording data to obtain voiceprint data;

[0025] The text conversion module is used to process trip recording data based on ASR technology to obtain voice-to-text data;

[0026] The scene recognition module is used to input trip recording data, voiceprint data, and voice text data into the serially running empty scene recognition model, turn signal scene recognition model, in-vehicle occupant scene recognition model, and navigation sound scene recognition model to obtain scene recognition results;

[0027] The scene tagging module is used to generate corresponding scene tags for ride-hailing orders that have the risk of fraudulent orders based on the scene recognition results;

[0028] The order fraud identification module is used to input the scene tags and ride-hailing order data corresponding to the order fraud identification model to identify order fraud, and output the corresponding order fraud risk tags based on the results of the order fraud identification.

[0029] As a preferred implementation, when the scene recognition module runs, it specifically performs the following:

[0030] The scene recognition model is activated by inputting trip recording data, voiceprint data, and speech-text data into multiple scene recognition models running in series. These models include an empty vehicle scene recognition model, a turn signal scene recognition model, an in-vehicle occupant scene recognition model, and a navigation sound scene recognition model. The empty vehicle scene recognition model includes several models for identifying various empty vehicle conditions, performing semantic analysis based on trip recording and speech-text data to determine the empty vehicle status. The turn signal scene recognition model includes several turn signal models for identifying different turn signal conditions, performing semantic analysis based on the voiceprint and voiceprint data of the turn signal prompt sounds. The system performs a line matching comparison to obtain the turn signal status; the in-vehicle occupancy scene recognition model includes several passenger number models, analyzes the number of human voiceprints in the trip recording data and voiceprint data, and performs semantic analysis on the speech text data to obtain the number of people in the vehicle; the navigation sound scene recognition model includes several navigation sound models, performs voiceprint analysis on the voiceprint data based on the navigation sound, and performs semantic analysis on the corresponding speech text data of the navigation sound to obtain the navigation status; after parallel processing of empty vehicle status, turn signal status, passenger number status, and navigation status, the corresponding scene recognition results are output.

[0031] As a preferred implementation, when the order-brushing identification module is running, it specifically performs the following:

[0032] After obtaining the scene tags generated from the scene recognition results of multiple serially running scene recognition models, the fraudulent order recognition model is run. The scene tags, as well as the voiceprint data and voice-text data corresponding to the ride-hailing order, are input into the fraudulent order recognition model. The fraudulent order recognition model performs fraudulent order recognition based on the voiceprint data and voice-text data. Combined with the ride-hailing order data, the scene tags are re-verified and corrected, and the corresponding fraudulent order risk tags are generated and output. The fraudulent order risk tags are then added to the corresponding ride-hailing order.

[0033] As a preferred embodiment, when the voiceprint analysis module is running, it specifically performs the following:

[0034] The voiceprints of drivers corresponding to ride-hailing orders are pre-registered in the voiceprint database. Trip recording data is input into the voiceprint model to analyze the number and type of voiceprints in the trip recordings. The voiceprints of turn signal prompts, navigation sounds, and human voices are extracted. The driver's voiceprint is removed from the human voices through voiceprint matching to obtain the passenger's voiceprint. The voiceprints of turn signal prompts, navigation sounds, and passenger voices are then integrated into voiceprint data.

[0035] In a preferred embodiment, when the text conversion module runs, it specifically performs the following:

[0036] The trip recording data is input into the ASR model, which converts the trip recording data into corresponding text to form speech-text data.

[0037] Compared with the prior art, the beneficial effects of the present invention are:

[0038] This invention combines trip recording data with voiceprint data and voice-text data converted from the trip recording data, and inputs them into a series of models for identifying empty rides, turn signals, passengers inside the vehicle, and navigation sounds. This multi-dimensional analysis of ride-hailing order fraud risks enables efficient and accurate identification of ride-hailing order fraud, reducing losses to ride-hailing platforms caused by fraudulent activities, lowering the false positive rate of fraud, and improving the driver experience. Attached Figure Description

[0039] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings, wherein:

[0040] Figure 1 This is a flowchart illustrating the method for screening empty order brushing based on voiceprint and ASR technology according to the present invention.

[0041] Figure 2 This is a schematic diagram of the device for screening empty order brushing based on voiceprint and ASR technology according to the present invention. Detailed Implementation

[0042] Preferred embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While preferred embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.

[0043] The term "comprising" and its variations as used herein signify open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "one example embodiment" and "one embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0044] Example 1

[0045] like Figure 1 As shown, in a first aspect, embodiments of the present invention disclose a method for screening fraudulent order placement based on voiceprint and ASR technology, specifically including the following steps:

[0046] Step S1: Obtain the trip recording data corresponding to the ride-hailing order.

[0047] Specifically, the server retrieves ride-hailing orders from the order database and then retrieves the corresponding trip recording data from the recording database. The trip recording data consists of in-vehicle sounds recorded by the driver or vehicle terminal during the ride-hailing order process.

[0048] Step S2: Perform voiceprint analysis on the trip recording data to obtain voiceprint data.

[0049] Specifically, the server pre-registers the voiceprints of the drivers corresponding to the ride-hailing orders into the voiceprint database, inputs the trip recording data into the voiceprint model, analyzes the number and type of voiceprints in the trip recordings, extracts the turn signal prompt voiceprints, navigation voiceprints and human voiceprints, removes the driver's voiceprint from the human voice voiceprints through voiceprint matching to obtain the passenger's voiceprint, and integrates the turn signal prompt voiceprints, navigation voiceprints and passenger voiceprints into voiceprint data.

[0050] Step S3: Process the trip recording data based on ASR technology to obtain voice text data.

[0051] Specifically, the server inputs the trip recording data into the ASR model, and the ASR model converts the trip recording data into corresponding text to form speech-text data.

[0052] Step S4: Input the trip recording data, voiceprint data, and voice text data into the serially running empty scene recognition model, turn signal scene recognition model, in-vehicle occupant scene recognition model, and navigation sound scene recognition model respectively, and obtain the scene recognition results.

[0053] Specifically, the server starts the scene recognition model and inputs the trip recording data, voiceprint data, and voice text data into multiple scene recognition models running in series, including the empty scene recognition model, turn signal scene recognition model, in-vehicle occupancy scene recognition model, and navigation sound scene recognition model.

[0054] The empty-load scenario recognition model includes several empty-load models for recognizing various empty-load situations. It performs semantic analysis based on trip recordings and voice text data to obtain the empty-load situation.

[0055] The turn signal scene recognition model includes several turn signal models for recognizing different turn signal conditions. It compares the matching degree of the voiceprint of the turn signal prompt sound with the voiceprint data to determine whether there is a mismatch between the turn signal and the driving path in order to obtain the turn signal condition.

[0056] The in-vehicle occupancy scene recognition model includes several passenger occupancy models, analyzes the number of human voiceprints in the trip recording data and voiceprint data, and obtains the occupancy information by analyzing the semantics of the speech and text data.

[0057] The navigation sound scene recognition model includes several navigation sound models. Based on the voiceprint of the navigation sound, voiceprint data is analyzed, and semantic analysis is performed on the corresponding speech-text data of the navigation sound to obtain the navigation information.

[0058] Finally, after parallel processing of the empty vehicle status, turn signal status, passenger count, and navigation status output by the empty vehicle scene recognition model, turn signal scene recognition model, in-vehicle passenger scene recognition model, and navigation sound scene recognition model, the corresponding scene recognition results are output. Furthermore, by combining the turn signal and navigation status, the degree of matching between the expected route of a ride-hailing order and the actual driving route can be determined. Ride-hailing orders where the turn signal and navigation do not match the expected route are filtered out. Further matching between the actual driving route of the ride-hailing vehicle and the expected route of the order is then performed to filter out cases of fraudulent orders.

[0059] Step S5: Based on the scene recognition results, generate corresponding scene tags for ride-hailing orders that have the risk of fraudulent orders.

[0060] Specifically, based on the scene recognition results, the server analyzes the specific risks of fraudulent orders in the scene recognition results, defines ride-hailing orders with at least one risk as ride-hailing orders with fraudulent orders, and generates corresponding scene tags for ride-hailing orders with fraudulent orders.

[0061] Step S6: Input the scene tag corresponding to the ride-hailing order and the ride-hailing order data into the order fraud identification model for order fraud identification, and output the corresponding order fraud risk tag based on the order fraud identification result.

[0062] Specifically, after the server obtains the scene tags generated from the scene recognition results of multiple serially running scene recognition models, it runs the order-brushing recognition model. The scene tags, as well as the voiceprint data and voice-text data corresponding to the ride-hailing order, are input into the order-brushing recognition model. The order-brushing recognition model performs order-brushing recognition based on the voiceprint data and voice-text data. Combined with the ride-hailing order data, the scene tags are re-verified and corrected, and the corresponding order-brushing risk tags are generated and output. The order-brushing risk tags are then added to the corresponding ride-hailing order.

[0063] Among them, ride-hailing orders with the risk of fraudulent orders will be entered into the fraudulent order risk database, where they can be further manually reviewed. The results of the manual review will be the final judgment, and drivers who commit fraudulent orders will be punished accordingly.

[0064] In summary, the working principle of the method for screening ride-hailing orders based on voiceprint and ASR technology described in this invention is as follows: By combining trip recording data with voiceprint data and voice text data converted from the trip recording data, the data is input into a series of interconnected empty-load scene recognition models, turn signal scene recognition models, in-vehicle passenger scene recognition models, and navigation sound scene recognition models. This multi-dimensional analysis of ride-hailing order fraud risks enables efficient and accurate identification of ride-hailing order fraud, thereby reducing the losses caused to ride-hailing platforms by fraudulent activities, lowering the false positive rate of fraudulent orders, and improving the driver experience.

[0065] Other steps of the method for screening empty-running order brushing based on voiceprint and ASR technology described in this embodiment are the same as those in the prior art.

[0066] Example 2

[0067] like Figure 2 As shown, in a second aspect, embodiments of the present invention disclose a device for screening fraudulent order placement based on voiceprint and ASR technology, comprising a recording acquisition module M1, a voiceprint analysis module M2, a text conversion module M3, a scene recognition module M4, a scene tagging module M5, and an order placement identification module M6, specifically including:

[0068] The audio recording acquisition module M1 is used to acquire trip audio recording data corresponding to ride-hailing orders;

[0069] The voiceprint analysis module M2 is used to perform voiceprint analysis on the travel recording data to obtain voiceprint data;

[0070] The text conversion module M3 is used to process trip recording data based on ASR technology to obtain voice-to-text data;

[0071] The scene recognition module M4 is used to input the trip recording data, voiceprint data and voice text data into the serially running empty scene recognition model, turn signal scene recognition model, in-vehicle occupant scene recognition model and navigation sound scene recognition model respectively, and obtain the scene recognition results;

[0072] The scene tagging module M5 is used to generate corresponding scene tags for ride-hailing orders with the risk of fraudulent orders based on scene recognition results;

[0073] The M6 ​​order-brushing identification module is used to input the scene tags and ride-hailing order data corresponding to the order-brushing identification model for order-brushing identification, and output the corresponding order-brushing risk tags based on the order-brushing identification results.

[0074] In a preferred embodiment of this example, when the scene recognition module M4 runs, it specifically performs the following: activating the scene recognition model, inputting the trip recording data, voiceprint data, and voice-text data into multiple scene recognition models running in series, specifically including an empty-load scene recognition model, a turn signal scene recognition model, an in-vehicle occupant scene recognition model, and a navigation sound scene recognition model. The empty-load scene recognition model includes several empty-load models for recognizing various empty-load conditions, performing semantic analysis based on the trip recording and voice-text data to obtain the empty-load status; the turn signal scene recognition model includes several turn signal models for recognizing different turn signal conditions. The system compares the voiceprint of the turn signal indicator with the voiceprint data to obtain the turn signal status. The in-vehicle occupant scene recognition model includes several passenger number models. It analyzes the number of human voiceprints in the trip recording data and voiceprint data, and performs semantic analysis on the speech text data to obtain the number of people in the vehicle. The navigation sound scene recognition model includes several navigation sound models. It performs voiceprint analysis on the voiceprint data based on the navigation sound and performs semantic analysis on the corresponding speech text data to obtain the navigation status. After processing the empty vehicle status, turn signal status, passenger number status, and navigation status in parallel, the corresponding scene recognition results are output.

[0075] Furthermore, when the order-brushing identification module M6 is running, it specifically performs the following: after obtaining the scene tags generated by the scene recognition results based on multiple serially running scene recognition models, it runs the order-brushing identification model, inputs the voiceprint data and voice text data corresponding to the ride-hailing order into the order-brushing identification model, performs order-brushing identification based on the voiceprint data and voice text data, re-verifies and corrects the scene tags in combination with the ride-hailing order data, generates and outputs the corresponding order-brushing risk tag, and adds the order-brushing risk tag to the corresponding ride-hailing order.

[0076] In a preferred embodiment, when the voiceprint analysis module M2 is running, it specifically performs the following: Pre-registering the voiceprints of drivers corresponding to ride-hailing orders in the voiceprint database; inputting trip recording data into the voiceprint model; analyzing the quantity and type of voiceprints in the trip recordings; extracting turn signal alert voiceprints, navigation voiceprints, and human voiceprints; removing driver voiceprints from the human voice voiceprints through voiceprint matching to obtain passenger voiceprints; and integrating the turn signal alert voiceprints, navigation voiceprints, and passenger voiceprints into voiceprint data. When the text conversion module M3 is running, it specifically performs the following: inputting trip recording data into the ASR model; the ASR model converts the trip recording data into corresponding text to form speech-text data.

[0077] In summary, the device for screening empty order brushing based on voiceprint and ASR technology described in this embodiment of the invention can implement all the steps of the method for screening empty order brushing based on voiceprint and ASR technology described in Embodiment 1 when it is running.

[0078] Other structures of the device for screening empty order brushing based on voiceprint and ASR technology described in this embodiment are referred to in the prior art.

[0079] Example 3

[0080] This invention also discloses an electronic device, comprising at least one processor and a memory communicatively connected to the at least one processor. The memory stores instructions executable by the at least one processor. When the at least one processor executes the instructions, it specifically implements the following steps: acquiring trip recording data corresponding to a ride-hailing order; performing voiceprint analysis on the trip recording data to acquire voiceprint data; processing the trip recording data based on ASR technology to acquire speech-text data; inputting the trip recording data, voiceprint data, and speech-text data into a series-running empty-load scene recognition model, a turn signal scene recognition model, a vehicle occupancy scene recognition model, and a navigation sound scene recognition model, respectively, to acquire scene recognition results; generating corresponding scene tags for ride-hailing orders with a risk of fraudulent order placement based on the scene recognition results; inputting the scene tags corresponding to the ride-hailing order and the ride-hailing order data into a fraudulent order identification model for fraudulent order identification, and outputting corresponding fraudulent order risk tags based on the fraudulent order identification results.

[0081] Example 4

[0082] This invention also discloses a storage medium storing a computer program. When the computer program is executed by a processor, it specifically implements the following steps: acquiring trip recording data corresponding to ride-hailing orders; performing voiceprint analysis on the trip recording data to acquire voiceprint data; processing the trip recording data based on ASR technology to acquire speech-text data; inputting the trip recording data, voiceprint data, and speech-text data into a series of interconnected empty-load scene recognition models, turn signal scene recognition models, in-vehicle passenger scene recognition models, and navigation sound scene recognition models to acquire scene recognition results; generating corresponding scene tags for ride-hailing orders with fraudulent order risks based on the scene recognition results; inputting the scene tags corresponding to the ride-hailing orders and the ride-hailing order data into a fraudulent order recognition model for fraudulent order recognition, and outputting corresponding fraudulent order risk tags based on the fraudulent order recognition results.

[0083] This disclosure can be a method, apparatus, system, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this disclosure.

[0084] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0085] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0086] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, Java, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0087] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0088] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0089] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0090] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which includes one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0091] Various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for screening fraudulent order placement based on voiceprint and ASR technology, characterized in that, Includes the following steps: Obtain the trip recording data corresponding to the ride-hailing order; Perform voiceprint analysis on the trip recording data to obtain voiceprint data; The trip recording data is processed using ASR technology to obtain voice-text data; The trip recording data, voiceprint data, and voice text data are respectively input into the cascaded empty scene recognition model, turn signal scene recognition model, in-vehicle occupant scene recognition model, and navigation sound scene recognition model to obtain scene recognition results. Based on the scene recognition results, corresponding scene tags are generated for ride-hailing orders that have the risk of fraudulent orders; Input the scene tags and ride-hailing order data corresponding to the ride-hailing order into the fraudulent order identification model to identify fraudulent orders, and output the corresponding fraudulent order risk tags based on the results of the fraudulent order identification. The process involves inputting the scene tags and ride-hailing order data corresponding to the ride-hailing order into the fraudulent order identification model for fraudulent order identification, and outputting corresponding fraudulent order risk tags based on the fraudulent order identification results. Specifically, this includes: After obtaining the scene tags generated from the scene recognition results of multiple serially running scene recognition models, the fraudulent order recognition model is run. The scene tags, as well as the voiceprint data and voice-text data corresponding to the ride-hailing order, are input into the fraudulent order recognition model. The fraudulent order recognition model performs fraudulent order recognition based on the voiceprint data and voice-text data. Combined with the ride-hailing order data, the scene tags are re-verified and corrected, and the corresponding fraudulent order risk tags are generated and output. The fraudulent order risk tags are then added to the corresponding ride-hailing order.

2. The method for screening fraudulent order processing based on voiceprint and ASR technology according to claim 1, characterized in that, The acquisition of scene recognition results specifically includes: The scene recognition model is activated by inputting trip recording data, voiceprint data, and speech-text data into multiple scene recognition models running in series. These models include an empty vehicle scene recognition model, a turn signal scene recognition model, an in-vehicle occupant scene recognition model, and a navigation sound scene recognition model. The empty vehicle scene recognition model includes several models for identifying various empty vehicle conditions, performing semantic analysis based on trip recording and speech-text data to determine the empty vehicle status. The turn signal scene recognition model includes several turn signal models for identifying different turn signal conditions, performing semantic analysis based on the voiceprint and voiceprint data of the turn signal prompt sounds. The system performs a line matching comparison to obtain the turn signal status; the in-vehicle occupancy scene recognition model includes several passenger number models, analyzes the number of human voiceprints in the trip recording data and voiceprint data, and performs semantic analysis on the speech text data to obtain the number of people in the vehicle; the navigation sound scene recognition model includes several navigation sound models, performs voiceprint analysis on the voiceprint data based on the navigation sound, and performs semantic analysis on the corresponding speech text data of the navigation sound to obtain the navigation status; after parallel processing of empty vehicle status, turn signal status, passenger number status, and navigation status, the corresponding scene recognition results are output.

3. The method for screening fraudulent order processing based on voiceprint and ASR technology according to claim 1, characterized in that, The process of performing voiceprint analysis on the trip recording data to obtain voiceprint data specifically includes: The voiceprints of drivers corresponding to ride-hailing orders are pre-registered in the voiceprint database. Trip recording data is input into the voiceprint model to analyze the number and type of voiceprints in the trip recordings. The voiceprints of turn signal prompts, navigation sounds, and human voices are extracted. The driver's voiceprint is removed from the human voices through voiceprint matching to obtain the passenger's voiceprint. The voiceprints of turn signal prompts, navigation sounds, and passenger voices are then integrated into voiceprint data.

4. The method for screening fraudulent order placement based on voiceprint and ASR technology according to claim 1, characterized in that, The process of processing trip recording data based on ASR technology to obtain voice-text data specifically includes: The trip recording data is input into the ASR model, which converts the trip recording data into corresponding text to form speech-text data.

5. A device for screening fraudulent order processing based on voiceprint and ASR technology, characterized in that, It includes a recording acquisition module, a voiceprint analysis module, a text conversion module, a scene recognition module, a scene tagging module, and a fraudulent order detection module, specifically including: The audio recording acquisition module is used to acquire trip audio recording data corresponding to ride-hailing orders; The voiceprint analysis module is used to perform voiceprint analysis on the travel recording data to obtain voiceprint data; The text conversion module is used to process trip recording data based on ASR technology to obtain voice-to-text data; The scene recognition module is used to input trip recording data, voiceprint data, and voice text data into the serially running empty scene recognition model, turn signal scene recognition model, in-vehicle occupant scene recognition model, and navigation sound scene recognition model to obtain scene recognition results; The scene tagging module is used to generate corresponding scene tags for ride-hailing orders that have the risk of fraudulent orders based on the scene recognition results; The order fraud identification module is used to input the scene tags and ride-hailing order data corresponding to the ride-hailing order into the order fraud identification model for order fraud identification, and output the corresponding order fraud risk tags based on the order fraud identification results; Specifically, when the order-brushing identification module is running, it performs the following: After obtaining the scene tags generated from the scene recognition results of multiple serially running scene recognition models, the fraudulent order recognition model is run. The scene tags, as well as the voiceprint data and voice-text data corresponding to the ride-hailing order, are input into the fraudulent order recognition model. The fraudulent order recognition model performs fraudulent order recognition based on the voiceprint data and voice-text data. Combined with the ride-hailing order data, the scene tags are re-verified and corrected, and the corresponding fraudulent order risk tags are generated and output. The fraudulent order risk tags are then added to the corresponding ride-hailing order.

6. The method for screening fraudulent order processing based on voiceprint and ASR technology according to claim 5, characterized in that, When the scene recognition module runs, it specifically performs the following: The scene recognition model is activated by inputting trip recording data, voiceprint data, and speech-text data into multiple scene recognition models running in series. These models include an empty vehicle scene recognition model, a turn signal scene recognition model, an in-vehicle occupant scene recognition model, and a navigation sound scene recognition model. The empty vehicle scene recognition model includes several models for identifying various empty vehicle conditions, performing semantic analysis based on trip recording and speech-text data to determine the empty vehicle status. The turn signal scene recognition model includes several turn signal models for identifying different turn signal conditions, performing semantic analysis based on the voiceprint and voiceprint data of the turn signal prompt sounds. The system performs a line matching comparison to obtain the turn signal status; the in-vehicle occupancy scene recognition model includes several passenger number models, analyzes the number of human voiceprints in the trip recording data and voiceprint data, and performs semantic analysis on the speech text data to obtain the number of people in the vehicle; the navigation sound scene recognition model includes several navigation sound models, performs voiceprint analysis on the voiceprint data based on the navigation sound, and performs semantic analysis on the corresponding speech text data of the navigation sound to obtain the navigation status; after parallel processing of empty vehicle status, turn signal status, passenger number status, and navigation status, the corresponding scene recognition results are output.

7. The device for screening empty order brushing based on voiceprint and ASR technology according to claim 5, characterized in that, When the voiceprint analysis module is running, it specifically performs the following: The voiceprints of drivers corresponding to ride-hailing orders are pre-registered in the voiceprint database. Trip recording data is input into the voiceprint model to analyze the number and type of voiceprints in the trip recordings. The voiceprints of turn signal prompts, navigation sounds, and human voices are extracted. The driver's voiceprint is removed from the human voices through voiceprint matching to obtain the passenger's voiceprint. The voiceprints of turn signal prompts, navigation sounds, and passenger voices are then integrated into voiceprint data.

8. The device for screening fraudulent order processing based on voiceprint and ASR technology according to claim 5, characterized in that, When the text conversion module runs, it specifically performs the following: The trip recording data is input into the ASR model, which converts the trip recording data into corresponding text to form speech-text data.

Citation Information

Patent Citations

  • Detection method and system for preventing driver from scalping based on vehicle-mounted system

    CN112926881A

  • Event identification method and device, equipment and storage medium

    CN113239872A