Order identification method and device, equipment and storage medium
By periodically sampling audio data and using text classification and order classification models to identify high-risk orders, the passive and lagging issues of high-risk order identification in ride-hailing scenarios are solved, achieving efficient and accurate safety warnings and interventions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies for identifying high-risk orders in ride-hailing scenarios are passive and delayed, making it difficult to provide effective early warnings and proactive interventions before danger occurs, resulting in high safety risks for drivers and passengers.
By sampling audio data at intervals, a text classification model is used to initially screen orders containing sexually sensitive words. Then, by combining the order classification model to extract feature vectors, high-risk orders are predicted, thus achieving accurate identification of high-risk orders.
It improved the accuracy of identifying high-risk orders, reduced the probability of related cases, and enabled safety warnings and proactive intervention in ride-hailing scenarios.
Smart Images

Figure CN121743973A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of order recognition technology, and in particular to an order recognition method, apparatus, device and storage medium. Background Technology
[0002] With the rapid development of the ride-hailing industry, people's travel options have become increasingly convenient. However, while enjoying this convenience, safety risks for both drivers and passengers have also emerged. Because ride-hailing scenarios involve unfamiliar drivers and passengers, enclosed spaces, and temporary environments, they objectively provide covert conditions for illegal activities, posing a real threat to the personal safety of some passengers, especially women.
[0003] However, current mainstream solutions rely on user-triggered features such as one-click alarms and trip sharing, which are often difficult for victims to perform in the tense environment of a closed vehicle. Existing technologies for identifying high-risk orders in ride-hailing scenarios are significantly passive and lagging. Although platforms possess trip trajectory data, they struggle to promptly capture potential safety hazards such as driver-passenger conflicts, sexual harassment, and assaults, resulting in an inability to provide effective warnings and proactive intervention before danger occurs. Summary of the Invention
[0004] This application provides an order identification method, apparatus, device, and storage medium. Based on intermittently sampled audio data, it uses a text classification model to initially screen the identified text, and then uses an order classification model to further identify high-risk orders based on trip recordings and order information. This improves the accuracy of identifying high-risk orders, thereby increasing the accuracy of judging the nature of orders in ride-hailing scenarios and reducing the probability of related cases.
[0005] In a first aspect, embodiments of this application propose an order identification method, including: During the order process, the trip recording is sampled at intervals, and the collected audio data is converted into first text with a preset first precision. When it is determined that the text contains preset sensitive words, the text is input into a text classification model to predict the order sensitivity probability. If the output first prediction probability is greater than the first preset threshold, the order information and trip recording of the order are obtained. The order information and trip recordings are input into the order classification model. The order classification model is used to extract the feature vectors of the order information and trip recordings respectively and concatenate them. The order risk probability is predicted based on the concatenated feature vectors. If the output second predicted probability is greater than the second preset threshold, then the order is determined to be a high-risk order.
[0006] In some possible embodiments, before obtaining the order information and trip recordings of the order, the method further includes: The text is input into the text classification model to predict the probability of order calls. If the output third prediction probability is less than the third preset threshold and the first prediction probability is greater than the first preset threshold, then the order information and trip recording of the order are obtained.
[0007] In some possible embodiments, the text classification model is trained through the following steps: The first training set is input into the text classification model. The first training set includes at least one text with a binary label, which is used to characterize whether the order corresponding to the text is a sexual order. The text classification model is used to predict the order sensitivity probability, and the parameters of the text classification model are adjusted with the goal of minimizing the loss function between the order sensitivity probability output by the text classification model and the binary label corresponding to each text. The test set is input into the adjusted text classification model, the accuracy of the text classification model is calculated, and the text classification model training ends when the accuracy is greater than a preset threshold.
[0008] In some possible embodiments, the first training set is obtained through the following steps: Obtain at least one text and label the first batch of texts to obtain the first training text; The first training text is labeled using a pre-trained labeling model, and the accuracy of the labeling model's labeling compared to the original labeling is calculated. The parameters of the labeling model are continuously adjusted so that the accuracy of the labeling model's labeling reaches a preset threshold. After determining that the accuracy of the labeling model in labeling the first training text reaches a preset threshold, the labeling model is used to label the remaining second batch of text to obtain the labeled second training text, wherein the first batch is smaller than the second batch; The first training text and the second training text are used as the first training set.
[0009] In some possible embodiments, before converting the acquired audio data into a first text with a preset first precision, the method further includes: The collected audio data is converted into second text with a preset second precision, wherein the second precision is lower than the first precision; When it is determined that the second text contains preset basic sensitive words, the collected audio data is converted into a first text with a preset first precision.
[0010] In some possible embodiments, the step of using the order classification model to extract feature vectors from the order information and trip recordings respectively and concatenating them, and predicting the order risk probability based on the concatenated feature vectors, includes: Using the feature extraction layer in the order classification model, at least one feature vector of the order information and the feature vector of the trip recording are extracted respectively. Using the feature concatenation layer in the order classification model, at least one feature vector of the order information is normalized and then concatenated with the feature vector of the trip recording; The order risk probability is predicted based on the concatenated feature vector by using the classifier prediction layer in the order classification model.
[0011] In some possible embodiments, the order classification model is trained through the following steps: The second training set is input into the order classification model. The second training set includes order information and trip recordings of at least one order with a second binary label. The second binary label is used to characterize whether the order is a high-risk order. The order classification model is used to predict the probability of order risk, and the parameters of the classifier prediction layer in the order classification model are adjusted with the goal of minimizing the loss function between the order risk probability output by the order classification model and the second binary label corresponding to each order.
[0012] Secondly, embodiments of this application provide an order identification device, comprising: The acquisition module is used to periodically sample the trip recordings during the order process and convert the acquired audio data into a first text with a preset first precision. The sensitive identification module is used to determine that when the text contains preset sensitive words, input the text into a text classification model to predict the order sensitivity probability. If the output first prediction probability is greater than a first preset threshold, the order information and trip recording of the order are obtained. The high-risk identification module is used to input the order information and trip recording into the order classification model. Using the order classification model, feature vectors of the order information and trip recording are extracted and concatenated. Based on the concatenated feature vector, the order risk probability is predicted. If the output second predicted probability is greater than the second preset threshold, the order is determined to be a high-risk order.
[0013] Thirdly, embodiments of this application provide an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform steps in an order identification method as described in any of the first aspects above.
[0014] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions for performing steps in an order identification method as described in any of the first aspects above.
[0015] The order identification method, apparatus, device, and storage medium provided in this application embodiment, based on intermittently sampled audio data, uses a text classification model to initially screen the identified text, and then uses an order classification model based on trip recordings and order information to further identify high-risk orders, thereby improving the accuracy of identifying high-risk orders, thus improving the accuracy of judging the nature of orders in ride-hailing scenarios and reducing the probability of related cases.
[0016] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating an order identification method according to an embodiment of this application; Figure 2 This is an exemplary flowchart of a text preprocessing-based embodiment of this application; Figure 3 This is a flowchart illustrating the training process of a text classification model in an embodiment of this application. Figure 4 This is a flowchart illustrating the annotation of a large number of training texts in an embodiment of this application; Figure 5 This is a schematic diagram of an optimized text classification model architecture in an embodiment of this application; Figure 6 This is a flowchart illustrating the training process of an order classification model according to an embodiment of this application. Figure 7 This is a flowchart illustrating a specific process for identifying an order to be identified in an embodiment of this application; Figure 8 This is a schematic diagram of an order recognition device according to an embodiment of this application; Figure 9 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0019] To further illustrate the technical solutions provided in the embodiments of this application, a detailed description is provided below in conjunction with the accompanying drawings and specific implementation methods.
[0020] With the rapid development of the ride-hailing industry, people's travel options have become increasingly convenient. However, while enjoying this convenience, safety risks for both drivers and passengers have also emerged. Because ride-hailing scenarios involve unfamiliar drivers and passengers, enclosed spaces, and temporary environments, they objectively provide covert conditions for illegal activities, posing a real threat to the personal safety of some passengers, especially women.
[0021] However, current mainstream solutions rely on user-triggered features such as one-click alarms and trip sharing, which are often difficult for victims to perform in the tense environment of a closed vehicle. Existing technologies for identifying high-risk orders in ride-hailing scenarios are significantly passive and lagging. Although platforms possess trip trajectory data, they struggle to promptly capture potential safety hazards such as driver-passenger conflicts, sexual harassment, and assaults, resulting in an inability to provide effective warnings and proactive intervention before danger occurs.
[0022] Given the aforementioned passivity and lag in high-risk order identification technologies, which prevent effective early warning and intervention before risks occur, this application proposes an order identification method, such as... Figure 1 As shown, it includes: Step S101: During the order process, the trip recording is sampled at intervals, and the collected audio data is converted into a first text with a preset first precision. Among them, existing ASR (Automatic Speech Recognition) technology is used to convert at least one audio data point into corresponding text. Step S102: When it is determined that the text contains preset sensitive words, the text is input into a text classification model to predict the order sensitivity probability. If the output first prediction probability is greater than the first preset threshold, the order information and trip recording of the order are obtained. The trip recording is the trip recording from the start of the order to the current moment, and the current moment is the moment when the text is determined to contain preset sexually sensitive words; Step S103: Input the order information and trip recording into the order classification model. Using the order classification model, extract the feature vectors of the order information and trip recording respectively and concatenate them. Based on the concatenated feature vector, predict the order risk probability. Step S104: If the output second predicted probability is greater than the second preset threshold, then the order is determined to be a high-risk order.
[0023] In this embodiment, order audio data obtained through intermittent real-time sampling is converted into corresponding text during the order process. Before using the order classification model to predict the order risk probability, preliminary screening is performed based on the text. Orders corresponding to texts containing preset sexually sensitive words are selected as candidate orders for further identification of high-risk orders. This improves the efficiency of identifying high-risk orders and provides more focused risk order samples for the order classification model, thereby enhancing the overall accuracy of high-risk order identification.
[0024] In some possible embodiments, the preset sexually sensitive words include, but are not limited to, specific words, phrases and common variations related to sexual behavior, sexual innuendo, invitations or transactions for sexual services. These are usually manually set and stored in a corresponding sexually sensitive word library. After obtaining the first text, based on the sexually sensitive word library, it is determined whether the first text includes preset sexually sensitive words. Specifically, if there is a word in the first text that is exactly the same as any preset sexually sensitive word, it is determined that the first text includes preset sexually sensitive words.
[0025] In this embodiment of the application, existing text matching methods or existing similarity detection algorithms can be used to determine whether the first text contains preset sexually sensitive words. No specific limitations are made in this embodiment of the application.
[0026] In some possible embodiments, to further reduce the interference of low-quality text data on the above-mentioned text classification model and improve processing efficiency, before converting the collected audio data into first text with a preset first precision, the method further includes: The collected audio data is converted into second text with a preset second precision, wherein the second precision is lower than the first precision; When it is determined that the second text contains preset basic sensitive words, the collected audio data is converted into a first text with a preset first precision.
[0027] In some possible embodiments, the accuracy represents the training accuracy of the corresponding speech recognition model. For example, the second-accuracy speech recognition model mentioned above is a speech recognition model trained with the goal of achieving an accuracy of 80% in matching the output text semantics with the standard audio semantics. The corresponding first-accuracy speech recognition model is trained with the goal of achieving an accuracy of 90% in matching the output text semantics with the standard audio semantics. In this application embodiment, only one possible implementation of the above-mentioned accuracy is given. In practical applications, existing speech recognition models can also be used, and the corresponding accuracy can be adjusted by setting the parameters of the model. No specific limitation is made here.
[0028] In this embodiment of the application, the preset basic sensitive words, including but not limited to, are usually manually set and stored in the corresponding basic sensitive word library. This is the same as the method described above for determining that the first text includes preset sensitive words. Existing text matching methods or existing similarity detection algorithms can be used. In this embodiment of the application, no specific limitations are made and will not be elaborated further.
[0029] Below is an exemplary workflow based on text preprocessing before using a text classification model for coarse screening, such as... Figure 2 As shown, it includes the following steps: Step S201: During the order process, the trip recording is sampled at intervals, and the collected audio data is converted into second text with a preset second precision. Step S202: Based on the preset basic sensitive word library, determine whether the second text contains preset basic sensitive words; Step S203a: When it is determined that the second text contains preset basic sensitive words, the collected audio data is converted into a first text with a preset first precision. Step S203b: If it is determined that the second text does not contain preset basic sensitive words, no intervention will be performed; Step S204: Based on a preset sex-related sensitive word library, determine whether the first text contains preset sex-related sensitive words; Step S205a: When it is determined that the first text contains preset sexually sensitive words, the text is input into a text classification model to predict the order sensitivity probability. Step S205b: If it is determined that the first text does not contain preset sexually sensitive words, no intervention is performed.
[0030] In this embodiment of the application, the text classification model is trained through the following steps, such as... Figure 3 As shown, it includes: Step S301: Input the first training set into the text classification model. The first training set includes at least one text with a binary label, which is used to characterize whether the order corresponding to the text is a sexual order. Optionally, a binary label of 1 indicates that the order corresponding to the text is a sexual order, and a binary label of 0 indicates that the order corresponding to the text is a non-sexual order. Step S302: Use the text classification model to predict the order sensitivity probability, and adjust the parameters of the text classification model with the goal of minimizing the loss function between the order sensitivity probability output by the text classification model and the binary label corresponding to each text. Step S303: Input the test set into the adjusted text classification model, calculate the accuracy of the text classification model, and end the text classification model training when the accuracy is greater than a preset threshold.
[0031] Specifically, the aforementioned text classification model can use the BERT-Based Chinese pre-trained model released by Google as the base model for text classification. Here, only one possible implementation is given. The model is trained on a first training set consisting of a large number of texts with binary labels, and tested on a test set to finally obtain a text classification model that achieves the required accuracy.
[0032] In some possible embodiments, the first training set mentioned above is obtained from orders corresponding to work orders related to sexually explicit scenarios pulled from the platform over a recent period (such as the past year). Specifically, all audio files of the order in progress are obtained and converted into text. The audio files are generally segmented, such as 30-second segments, 60-second segments, etc. Then, based on sexually explicit sensitive words defined in advance by humans, the converted text of the audio files is filtered to obtain text containing sexually explicit sensitive words, which is used as the training text in the first training set mentioned above.
[0033] Optionally, sensitive words can be extracted from the above text through the design of prompt words in other large models, and finally, sexually sensitive words can be obtained through manual review.
[0034] In this embodiment of the application, the binary label used to characterize whether the order corresponding to the text is a sexually explicit order can be manually labeled and / or labeled using a large model. In order to improve labeling efficiency and reduce labor costs, a large model labeling method can be used to label a large number of texts. By labeling each training text in the first training set, it can be determined whether each training text is sexually explicit text.
[0035] Optionally, the above-mentioned large-scale model annotation method is used to annotate a large amount of training text, such as... Figure 4 As shown, the first training set is obtained through the following steps: Step S401: Obtain at least one text and label the first batch of texts to obtain the first training text; Step S402: The first training text is labeled using a pre-trained labeling model, and the accuracy of the labeling model's labeling compared with the original labeling is calculated. The parameters of the labeling model are continuously adjusted so that the accuracy of the labeling model's labeling reaches a preset threshold. Step S403: After determining that the accuracy of the labeling model in labeling the first training text reaches a preset threshold, the labeling model is used to label the remaining second batch of text to obtain the labeled second training text, wherein the first batch is smaller than the second batch. Step S404: Use the first training text and the second training text as the first training set.
[0036] In some possible embodiments, after the large model is labeled, the results of the large model labeling can be manually sampled and checked to ensure the accuracy of the training text labeling. The large model used for text labeling can be any existing labelable large model; no specific limitation is made in this embodiment.
[0037] In some possible embodiments, step S303 of the text classification model training method further includes re-labeling the low-confidence texts predicted by the model (such as texts with a prediction probability between 0.3 and 0.7), adding the re-labeled texts to the first training set as training texts, and training the text classification model again until the accuracy of the test set does not improve significantly, that is, the accuracy improvement value is less than a preset threshold, and the training of the text classification model is completed.
[0038] In some possible embodiments, when actually using a text classification model to predict the order sensitivity probability, there may be errors such as misidentifying orders involving call scenarios as sexually sensitive orders. Therefore, in order to avoid the above errors as much as possible, in this embodiment of the application, before obtaining the order information and trip recording of the order, the method further includes: The text is input into the text classification model to predict the probability of order calls. If the output third prediction probability is less than the third preset threshold and the first prediction probability is greater than the first preset threshold, then the order information and trip recording of the order are obtained.
[0039] Optionally, because ride-hailing orders often involve adding each other on WeChat or making offline appointments during phone calls (whether by drivers or passengers), and these situations are similar to sexually suggestive text requests for contact information, the aforementioned text classification model is prone to false positives. Therefore, the existing model can be modified by adding a corresponding classification head to the outermost layer, causing the model to output two probability results: one for order sensitivity and one for order call probability.
[0040] In practical applications, we set two probability thresholds: one for the order's sensitivity probability and the other for the order's call probability. Only when the probabilities output by the model's two classification heads meet the preset conditions will the order be considered a sensitive order. The optimized text classification model architecture is as follows: Figure 5 As shown, it includes: The system consists of a pre-trained base model, a trained fully connected layer, and two independent classification heads. The trained fully connected layer is used to efficiently and specifically transform the general semantic representation generated by the pre-trained base model into professional features applicable to the current specific text classification task. The first classification head is used to output the prediction of order sensitivity probability; The second classification head is used to output the predicted probability of order calls.
[0041] In some possible embodiments, the text classification model architecture also includes a Dropout layer to prevent overfitting by randomly discarding some neurons.
[0042] In some possible embodiments, the preset condition is that the third predicted probability of the output order call probability is less than a third preset threshold, and the first predicted probability of the output order sensitivity probability is greater than a first preset threshold.
[0043] Correspondingly, the optimized text classification model is trained using the same training method described above. Specifically, a first training set is input into the text classification model. The first training set includes at least one text with a first binary label representing whether the order corresponding to the text is a sensitive order and a third binary label representing whether the order corresponding to the text is a call order. The text classification model is used to predict order sensitivity and call probability respectively. The parameters of the text classification model are adjusted with the goal of minimizing the loss function between the order sensitivity probability output by the text classification model and the first binary label corresponding to each text, and the loss function between the order call probability and the third binary label corresponding to each text. The test set is input into the adjusted text classification model, the accuracy of the text classification model is calculated, and the text classification model training ends when the accuracy is greater than a preset threshold.
[0044] Optionally, if the first binary label is 1, it means that the order corresponding to the text is a sexual order; if the first binary label is 0, it means that the order corresponding to the text is a non-sexual order. If the third binary label is 1, it means that the order corresponding to the text is a call order; if the third binary label is 0, it means that the order corresponding to the text is a non-call order.
[0045] In step S102 above, when using the trained text classification model to predict the order sensitivity probability of the current text containing preset sexually sensitive words, if the output first prediction probability is greater than a first preset threshold, then the order information and trip recording of the order are obtained, wherein the order information of the order includes at least one of the following: Call time period, mileage information, drop-off time, origin and destination, gender of driver and passenger.
[0046] In some possible embodiments, if the output first predicted probability is greater than a first preset threshold, the order will be automatically determined to have a basic sexual risk. After this determination is triggered, a preset primary safety intervention process can be automatically executed without waiting for manual review. The primary safety intervention process includes, but is not limited to, automatically triggering and locking a safety prompt pop-up window, contacting passengers via intelligent voice outbound calls, and other primary intervention processes.
[0047] Optionally, the content of the aforementioned intelligent voice outbound calls can be pre-prepared, non-intrusive, caring scripts (e.g., "Hello, the platform has detected that this trip is about to begin / is in progress. Please be careful. If you need assistance, please contact us through the security center in the app"). This not only sends a signal to passengers that they are being cared for, but also collects additional auxiliary information for risk assessment through recording and real-time passenger reactions.
[0048] In some possible embodiments, in step S103 above, the order classification model is used to extract feature vectors from the order information and the trip recording, and these feature vectors are then concatenated. Based on the concatenated feature vectors, the order risk probability is predicted, including: Using the feature extraction layer in the order classification model, at least one feature vector of the order information and the feature vector of the trip recording are extracted respectively. Using the feature concatenation layer in the order classification model, at least one feature vector of the order information is normalized and then concatenated with the feature vector of the trip recording; The order risk probability is predicted based on the concatenated feature vector by using the classifier prediction layer in the order classification model.
[0049] Specifically, at least one feature vector is extracted from the order information. Typically, three core methods are used depending on the data type: For categorical / classified features (such as call time, driver and passenger gender), one-hot encoding or embedding learning is commonly used to convert them into binary vectors or low-dimensional dense vectors; for numerical / continuous features (such as delivery time, racking time), standardization / normalization or binning discretization are used to preserve numerical relationships or convert them into categories; for spatial / geographical features (such as the origin and destination type of the order), direct coordinate encoding or geographic semantic encoding (such as mapping to regional type) is used to characterize its physical or business attributes.
[0050] In this embodiment, the order classification model utilizes a feature extraction layer to extract features from the order's call time period, mileage, delivery time, passenger gender, and origin / destination type, obtaining corresponding feature vectors. The Hubert deep learning model is then used to extract audio features from the trip recordings, extracting the output of the last hidden layer as the audio embedding vector. Next, the feature concatenation layer in the order classification model embeds the feature vectors extracted from the order information. Numerical features (e.g., order mileage, delivery time) are normalized and concatenated with the audio embedding vector to obtain a concatenated feature vector containing rich semantic and acoustic information. Finally, the classifier prediction layer in the order classification model predicts the order risk probability based on the concatenated feature vector.
[0051] In this embodiment of the application, the order classification model is trained through the following steps, such as... Figure 6 As shown, it includes: Step S601: Input the second training set into the order classification model. The second training set includes order information and trip recordings of at least one order with a second binary label. The second binary label is used to characterize whether the order is a high-risk order. Step S602: Predict the order risk probability using the order classification model, and adjust the parameters of the classifier prediction layer in the order classification model with the goal of minimizing the loss function between the order risk probability output by the order classification model and the second binary label corresponding to each order.
[0052] Optionally, the second training set required by the above order classification model comes from previous pop-up work orders. Specifically, based on the manual review results of the pop-up work orders, the order information and trip recordings of the corresponding orders are labeled with corresponding second binary labels. The second binary label is 1, which means that the order corresponding to the text is a high-risk order, and the second binary label is 0, which means that the order corresponding to the text is a non-high-risk order.
[0053] In some possible embodiments, if the output second predicted probability is greater than the second preset threshold, after determining that the order is a high-risk order, the method further includes triggering a corresponding preset manual intervention process, which includes, but is not limited to: contacting the driver and passenger simultaneously through a two-way safety call to verify and warn them in real time, or taking multi-dimensional intervention measures such as remotely controlling the trip (such as suspending the order or changing the destination to a safe area) based on the risk level judged by the human, and triggering in-vehicle device alarms or one-click notification of the passenger's emergency contact.
[0054] In addition, for extremely high-risk situations, it supports directly pushing key evidence packages and real-time location data to the police platform, initiating police-enterprise collaboration and recording the entire process to ensure compliant and traceable handling, ultimately forming a complete closed loop from real-time intervention to post-event safety follow-up.
[0055] Through the aforementioned primary security intervention process and manual intervention process, a tiered response mechanism from automatic early warning to in-depth manual intervention is achieved. This mechanism provides rapid response, a wide range of intervention methods, and system-wide linkage with external security resources, forming a closed loop for security handling. It also ensures human supervision and control at critical decision-making stages, achieving a dual guarantee of operational efficiency and security.
[0056] The following is a specific process for identifying orders using the above order identification method, wherein the text classification model is an optimized text classification model, such as... Figure 7 As shown, it includes the following steps: Step S701: During the order process, the trip recording is sampled at intervals, and the collected audio data is converted into second text with a preset second precision. Step S702: When it is determined that the second text contains preset basic sensitive words, the collected audio data is converted into a first text with a preset first precision. Step S703: When it is determined that the first text contains preset sexually sensitive words, the first text is input into the text classification model to predict the order sensitivity probability. Step S704: If the output first predicted probability is greater than the first preset threshold, then obtain the order information and trip recording of the order. Step S705: Input the order information and trip recording into the order classification model. Using the order classification model, extract the feature vectors of the order information and trip recording respectively and concatenate them. Based on the concatenated feature vector, predict the order risk probability. Step S706: If the output second predicted probability is greater than the second preset threshold, then the order is determined to be a high-risk order.
[0057] In this embodiment, a cascaded architecture employing a text classification model for coarse screening and an order classification model for fine screening significantly improves recognition accuracy while ensuring processing efficiency. Furthermore, it innovatively integrates structured order information with unstructured travel recordings, constructing a comprehensive chain of evidence. In addition, an iterative training mechanism combining manual annotation and large-scale model annotation enables continuous self-evolution of the text classification model's capabilities. This order recognition method not only significantly reduces manual review costs but also effectively reduces false positives through multi-dimensional cross-validation, providing a scalable and evolvable intelligent solution for platform security governance.
[0058] Based on the same inventive concept, embodiments of this application propose an order identification device, such as... Figure 8 As shown, it includes: The acquisition module 801 is used to periodically sample the trip recording during the order process and convert the acquired audio data into a first text with a preset first precision. The sensitive identification module 802 is used to determine that the text contains preset sensitive words, input the text into a text classification model to predict the order sensitivity probability, and if the output first prediction probability is greater than a first preset threshold, then obtain the order information and trip recording of the order. The high-risk identification module 803 is used to input the order information and trip recording into the order classification model, and use the order classification model to extract the feature vectors of the order information and trip recording respectively and concatenate them. Based on the concatenated feature vector, the order risk probability is predicted; if the output second predicted probability is greater than the second preset threshold, the order is determined to be a high-risk order.
[0059] The order recognition device described in this application embodiment is based on a cascaded recognition architecture that uses a text classification model for coarse screening and an order classification model for fine screening. The text classification model utilizes a small amount of manually labeled training text, combined with a large-scale model annotation method, and iteratively trains based on a pre-trained model, achieving high-quality, low-cost, and rapid initial screening. Furthermore, the order classification model integrates structured order information with unstructured travel recordings and other multimodal data, performing deep cross-validation and accurate discrimination. This significantly improves the accuracy and reliability of recognition while ensuring system processing efficiency, enabling efficient and accurate pre-screening of high-risk orders.
[0060] Based on the same inventive concept, embodiments of this application propose an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform an order identification method as described in any of the first aspects of the above embodiments.
[0061] The following reference Figure 9 This application describes an electronic device 900 according to one embodiment of the present application. Figure 9 The device 900 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0062] like Figure 9 As shown, an electronic device 900 is presented in the form of a general electronic device. The components of an electronic device 900 may include, but are not limited to: at least one processor 901, at least one memory 902, and a bus 903 connecting different system components (including memory 9002 and processor 9001).
[0063] Bus 903 represents one or more of several bus architectures, including a memory bus or memory controller, peripheral bus, processor, or local bus using any of the various bus architectures.
[0064] The memory 902 may include a readable medium in the form of volatile memory, such as random access memory (RAM) 9021 and / or cache memory 9022, and may further include read-only memory (ROM) 9023.
[0065] The memory 902 may also include a program / utility 9025 having a set (at least one) of program modules 9024, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0066] An electronic device 900 can also communicate with one or more external devices 904 (e.g., keyboard, pointing device, etc.), and with one or more devices that enable a user to interact with the electronic device 900, and / or with any device that enables the electronic device 900 to communicate with one or more other electronic devices (e.g., router, modem, etc.). This communication can be performed via an input / output (I / O) interface 905. Furthermore, the electronic device 900 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via a network adapter 906. As shown, the network adapter 906 communicates with other modules used in the electronic device 900 via a bus 903. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with the electronic device 900, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0067] Based on the same inventive concept, embodiments of this application provide a computer-readable storage medium storing a computer program. The computer program includes program instructions, which, when executed by a computer, cause the computer to perform any of the order identification methods discussed above. Since the principle by which the above-described computer-readable storage medium solves the problem is similar to that of the order identification method, the implementation of the above-described computer-readable storage medium can be referred to the implementation of the method; repeated details will not be elaborated further.
[0068] The order identification method, apparatus, device, and computer-readable storage medium described in this application embodiment, based on intermittently sampled audio data, uses a text classification model to initially screen the identified text, and then uses an order classification model based on trip recordings and order information to further identify high-risk orders, thereby improving the accuracy of identifying high-risk orders, thus improving the accuracy of judging the nature of orders in ride-hailing scenarios and reducing the probability of related cases occurring.
[0069] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0070] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0071] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0072] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0073] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. An order identification method, characterized in that, include: During the order process, the trip recording is sampled at intervals, and the collected audio data is converted into first text with a preset first precision. When it is determined that the text contains preset sensitive words, the text is input into a text classification model to predict the order sensitivity probability. If the output first prediction probability is greater than the first preset threshold, the order information and trip recording of the order are obtained. The order information and trip recordings are input into the order classification model. The order classification model is used to extract the feature vectors of the order information and trip recordings respectively and concatenate them. The order risk probability is predicted based on the concatenated feature vectors. If the output second predicted probability is greater than the second preset threshold, then the order is determined to be a high-risk order.
2. The method according to claim 1, characterized in that, Before obtaining the order information and trip recordings for the order, the method further includes: The text is input into the text classification model to predict the probability of order calls. If the output third prediction probability is less than the third preset threshold and the first prediction probability is greater than the first preset threshold, then the order information and trip recording of the order are obtained.
3. The method according to claim 1, characterized in that, The text classification model is trained through the following steps: The first training set is input into the text classification model. The first training set includes at least one training text with a first binary label. The first binary label is used to characterize whether the order corresponding to the training text is a sexual order. The text classification model is used to predict the order sensitivity probability, and the parameters of the text classification model are adjusted with the goal of minimizing the loss function between the order sensitivity probability output by the text classification model and the first binary label corresponding to each training text. The test set is input into the adjusted text classification model, the accuracy of the text classification model is calculated, and the text classification model training ends when the accuracy is greater than a preset threshold.
4. The method according to claim 3, characterized in that, The first training set was obtained through the following steps: Obtain at least one text and label the first batch of texts to obtain the first training text; The first training text is labeled using a pre-trained labeling model, and the accuracy of the labeling model's labeling compared to the original labeling is calculated. The parameters of the labeling model are continuously adjusted so that the accuracy of the labeling model's labeling reaches a preset threshold. After determining that the accuracy of the labeling model in labeling the first training text reaches a preset threshold, the labeling model is used to label the remaining second batch of text to obtain the labeled second training text, wherein the first batch is smaller than the second batch; The first training text and the second training text are used as the first training set.
5. The method according to claim 1, characterized in that, Before converting the acquired audio data into first text with a preset first precision, the method further includes: The collected audio data is converted into second text with a preset second precision, wherein the second precision is lower than the first precision; When it is determined that the second text contains preset basic sensitive words, the collected audio data is converted into a first text with a preset first precision.
6. The method according to claim 1, characterized in that, The step of using the order classification model to extract feature vectors from the order information and trip recordings, concatenating them, and predicting the order risk probability based on the concatenated feature vectors includes: Using the feature extraction layer in the order classification model, at least one feature vector of the order information and the feature vector of the trip recording are extracted respectively. Using the feature concatenation layer in the order classification model, at least one feature vector of the order information is normalized and then concatenated with the feature vector of the trip recording; The order risk probability is predicted based on the concatenated feature vector by using the classifier prediction layer in the order classification model.
7. The method according to claim 6, characterized in that, The order classification model is trained through the following steps: The second training set is input into the order classification model. The second training set includes order information and trip recordings of at least one order with a second binary label. The second binary label is used to characterize whether the order is a high-risk order. The order classification model is used to predict the probability of order risk, and the parameters of the classifier prediction layer in the order classification model are adjusted with the goal of minimizing the loss function between the order risk probability output by the order classification model and the second binary label corresponding to each order.
8. An order identification device, characterized in that, include: The acquisition module is used to periodically sample the trip recordings during the order process and convert the acquired audio data into a first text with a preset first precision. The sensitive identification module is used to determine that when the text contains preset sensitive words, input the text into a text classification model to predict the order sensitivity probability. If the output first prediction probability is greater than a first preset threshold, the order information and trip recording of the order are obtained. The high-risk identification module is used to input the order information and trip recording into the order classification model. Using the order classification model, feature vectors of the order information and trip recording are extracted and concatenated. Based on the concatenated feature vector, the order risk probability is predicted. If the output second predicted probability is greater than the second preset threshold, the order is determined to be a high-risk order.
9. An electronic device, characterized in that, include: At least one processor; The at least one processor is also connected in communication with a memory, wherein the memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the steps of an order identification method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for performing steps in an order identification method as described in any one of claims 1 to 7.