Data responsibility judgment processing method and device, equipment and medium

By obtaining and processing multimodal data in moving services, using feature extraction and multimodal large model for feature refining, and combining quantitative large model for responsibility analysis, the problems of inefficient quality inspection of violations of moving services in the existing technology and inaccurate judgment and result in inaccurate judgment and responsibility treatment are achieved efficient and accurate judgment and responsibility treatment.

CN119991356APending Publication Date: 2025-05-13SHENZHEN YISHIHUOLALA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510086724.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing quality inspection methods for moving service violations are inefficient and it is difficult to ensure the accuracy of the judgment results.

Method used

By obtaining multimodal data of moving personnel in moving services, including item packaging pictures, customer complaint text, order evaluation text, application voice information and GPS trajectory information, preset feature extraction and refinement using preset feature extraction strategies and multimodal models, and finally conducting judgment and analysis based on the quantitative model.

Benefits of technology

It has achieved efficient and accurate judgment and handling of the behavior of moving service personnel, improved the efficiency of judgment and handling of moving service behavior, and ensured the accuracy of the judgment and resulted in the judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991356A_ABST
    Figure CN119991356A_ABST
Patent Text Reader

Abstract

The invention provides a data responsibility judgment processing method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring multi-modal data of a moving person in a moving service; wherein the multi-modal data at least comprises an article packaging picture, a customer complaint text, an order evaluation text, application voice information and GPS track information; performing feature extraction on the multi-modal data based on a preset feature extraction strategy to obtain corresponding multi-modal features; performing feature extraction processing on the multi-modal features based on a preset multi-modal large model to obtain corresponding effective information; performing responsibility judgment analysis processing on the effective information based on a preset quantitative large model to obtain a responsibility judgment result corresponding to the moving personnel; and carrying out output processing on the responsibility judgment result. According to the invention, comprehensive, accurate and efficient responsibility judgment processing of the moving service behavior of the moving personnel is realized, the responsibility judgment processing efficiency of the moving service behavior is improved, and the accuracy of the responsibility judgment result is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a data judgment and processing method, device, computer equipment and storage medium. Background Art

[0002] With the booming housing rental market and the significant increase in urban population mobility in recent years, the demand for moving services has continued to grow. Against this backdrop, customers are increasingly demanding the quality of moving services, no longer just satisfied with the traditional speed of moving, but more focused on the safety and standardization of the service process, as well as the protection of personal privacy and property safety. In order to meet these increasingly diverse needs, many transportation companies have increased their control over moving service personnel, striving to comprehensively improve the overall level of moving services through a series of management measures.

[0003] In the current moving service management and control system, the main focus is on the following aspects: First, in terms of process, it is necessary to strictly supervise the service behavior of movers, including but not limited to checking whether there are violations such as "improper packing", "maliciously guiding customers to modify packages", "guiding customers to cancel orders", "obtaining customer contact information" and "asking about customer privacy"; secondly, in terms of service, focus on the service attitude of movers to ensure that they do not have problems such as "negative service attitude" and "bad service attitude"; finally, on the platform red line, severe crackdowns will be taken on serious violations such as "private transactions" by movers.

[0004] However, the existing quality inspection methods for moving service violations have many pain points. Specifically, due to the huge amount of data generated during the moving service process, including pictures, text, voice and other types of information, the traditional manual quality inspection method is not only time-consuming and labor-intensive, but also inefficient, and it is difficult to ensure the accuracy of the judgment results. This manual quality inspection model is often unable to cope with large-scale data and cannot meet the needs of efficient and accurate quality inspection, thus affecting the improvement of the overall quality of moving services and the improvement of customer satisfaction.

[0005] Therefore, in view of the various problems existing in the existing quality inspection methods for violations of moving services, it is urgent to develop a more efficient and accurate quality inspection technology to achieve real-time monitoring and precise judgment of the behavior of moving service personnel, thereby comprehensively improving the quality of moving services and customer satisfaction. Summary of the invention

[0006] The main purpose of the present invention is to provide a data accountability processing method, device, computer equipment and storage medium, aiming to solve the technical problems that the existing quality inspection methods for moving service violations have low processing efficiency and difficulty in ensuring the accuracy of the accountability results.

[0007] To achieve the above object, the present invention provides a data judgment processing method, which comprises:

[0008] Acquire multimodal data of movers in the moving service; wherein the multimodal data at least includes pictures of item packaging, customer complaint text, order evaluation text, application voice information and GPS track information;

[0009] Extracting features from the multimodal data based on a preset feature extraction strategy to obtain corresponding multimodal features;

[0010] Based on a preset multimodal large model, feature extraction processing is performed on the multimodal features to obtain corresponding effective information;

[0011] Based on a preset quantitative large model, the effective information is analyzed and processed to obtain a judgment result corresponding to the moving personnel;

[0012] The judgment result is outputted.

[0013] Optionally, the extracting features from the multimodal data based on a preset feature extraction strategy to obtain corresponding multimodal features includes:

[0014] Performing image feature extraction and text feature extraction on the item packaging pictures respectively to obtain corresponding matrix features and text vector features;

[0015] Performing text feature extraction on the customer complaint text to obtain corresponding text features;

[0016] Extracting emotional features from the order evaluation text to obtain a corresponding first emotional feature;

[0017] Performing voice feature extraction and emotion feature extraction on the application voice information to obtain corresponding voice features and second emotion features;

[0018] Extracting numerical features from the GPS trajectory information to obtain corresponding numerical features;

[0019] The matrix features, the text vector features, the text features, the first emotion features, the voice features, the second emotion features and the numerical features are integrated to obtain the multimodal features.

[0020] Optionally, the performing feature extraction processing on the multimodal features based on the preset multimodal macro model to obtain corresponding valid information includes:

[0021] Performing fusion processing on the multimodal features to obtain corresponding multimodal feature vectors;

[0022] Inputting the multimodal feature vector into a characterization capability layer in the multimodal large model, performing characterization processing on the multimodal feature vector through the characterization capability layer to obtain a corresponding first feature;

[0023] Inputting the first feature into the attention application layer in the multimodal large model, performing feature selection processing on the first feature through the attention application layer to obtain a corresponding second feature;

[0024] Decoding the second feature based on the decoder in the multimodal large model to obtain a corresponding output sequence;

[0025] The output sequence is used as the effective information.

[0026] Optionally, the fusing the multimodal features to obtain a corresponding multimodal feature vector includes:

[0027] Get the preset fusion strategy;

[0028] Based on the fusion strategy, the multimodal features are fused to obtain corresponding fusion features;

[0029] The fused features are used as the multimodal feature vector.

[0030] Optionally, the step of inputting the first feature into an attention application layer in the multimodal large model, performing feature selection processing on the first feature through the attention application layer to obtain a corresponding second feature includes:

[0031] Inputting the first feature into an attention application layer in the multimodal large model;

[0032] Performing weight allocation processing on the first feature based on the attention application layer to obtain an output importance weight;

[0033] Screening and listing the first features based on the importance weights to obtain corresponding initial features;

[0034] Combining the initial features to obtain a combined feature set;

[0035] The feature set is used as the second feature.

[0036] Optionally, the effective information is analyzed and processed based on a preset quantitative large model to obtain a judgment result corresponding to the mover, including:

[0037] Performing information analysis on the effective information based on the quantitative big model to extract key information related to the judgment;

[0038] Convert the key information into corresponding prompt words;

[0039] Sorting out the prompt words to obtain corresponding target prompt words;

[0040] Invoke the preset judgment rules;

[0041] Based on the judgment rule, the target prompt word is subjected to judgment analysis and processing to obtain a judgment result corresponding to the mover.

[0042] Optionally, the obtaining of multimodal data of movers in the moving service includes:

[0043] Acquiring initial multimodal data of the mover in the moving service;

[0044] Acquire a preprocessing strategy corresponding to the initial multimodal data;

[0045] Preprocessing the initial multimodal data based on the preprocessing strategy to obtain corresponding designated multimodal data;

[0046] The designated multimodal data is used as the multimodal data.

[0047] In addition, to achieve the above-mentioned purpose, the present invention further provides a data judgment processing device, the data judgment processing device comprising:

[0048] An acquisition module is used to acquire multimodal data of movers in the moving service; wherein the multimodal data at least includes pictures of item packaging, customer complaint text, order evaluation text, application voice information and GPS track information;

[0049] An extraction module, used to extract features from the multimodal data based on a preset feature extraction strategy to obtain corresponding multimodal features;

[0050] A first processing module is used to perform feature extraction processing on the multimodal features based on a preset multimodal large model to obtain corresponding valid information;

[0051] The second processing module is used to perform accountability analysis on the effective information based on a preset quantitative large model to obtain an accountability result corresponding to the mover;

[0052] The output module is used to output the judgment result.

[0053] In order to solve the above technical problems, the embodiment of the present application further provides a computer device, which adopts the following technical solution:

[0054] The computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of any one of the data judgment processing methods proposed in the embodiments of the present application are implemented.

[0055] In order to solve the above technical problems, the embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:

[0056] The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps of any one of the data judgment processing methods proposed in the embodiments of the present application are implemented.

[0057] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0058] The present invention provides a data accountability processing method, device, computer equipment and storage medium, the method comprising: firstly acquiring multimodal data of movers in moving services; wherein the multimodal data at least includes item packaging pictures, customer complaint texts, order evaluation texts, application voice information and GPS track information; then extracting features from the multimodal data based on a preset feature extraction strategy to obtain corresponding multimodal features; then performing feature refinement processing on the multimodal features based on a preset multimodal large model to obtain corresponding valid information; subsequently performing accountability analysis processing on the valid information based on a preset quantitative large model to obtain accountability results corresponding to the movers; and finally outputting the accountability results. The present invention obtains multimodal data of movers in moving services, then extracts features from the multimodal data based on a preset feature extraction strategy to obtain corresponding multimodal features, and then refines the multimodal features based on a multimodal big model to obtain corresponding effective information, and subsequently performs accountability analysis on the effective information based on a preset quantitative big model to obtain accountability results corresponding to the movers, and outputs the accountability results, thereby achieving comprehensive, accurate and efficient accountability processing for the moving service behaviors of the movers, improving the accountability processing efficiency for the moving service behaviors, and ensuring the accuracy of the obtained accountability results. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the scheme in the present application, a brief introduction is given below to the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0060] Figure 1is an exemplary system architecture diagram to which the present application may be applied;

[0061] Figure 2 is a flow chart of a data judgment processing method provided by an embodiment of the present invention;

[0062] Figure 3 is a structural schematic diagram of an embodiment of a data judgment and processing device according to the present application;

[0063] Figure 4 This is a basic structural block diagram of the computer device in this embodiment. DETAILED DESCRIPTION

[0064] The data judgment processing method provided by the embodiment of the present invention is applied to the data judgment processing device. Unless otherwise defined, all technical and scientific terms used in this document have the same meaning as those generally understood by technicians in the technical field of this application; the terms used in the specification of the application herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" in the specification and claims of this application and the above-mentioned drawings and any variations thereof are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.

[0065] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0066] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.

[0067] like Figure 1 As shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0068] Users can use terminal devices 101, 102, 103 to interact with server 105 through network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social online platform software, etc.

[0069] Terminal devices 101, 102, 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smart phones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers, desktop computers, etc.

[0070] The server 105 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal devices 101 , 102 , and 103 .

[0071] It should be noted that the data judgment and processing method provided in the embodiment of the present application is generally executed by a server / terminal device, and accordingly, the data judgment and processing device is generally arranged in the server / terminal device.

[0072] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0073] With the booming housing rental market and the significant increase in urban population mobility in recent years, the demand for moving services has continued to grow. Against this backdrop, customers are increasingly demanding the quality of moving services, no longer just satisfied with the traditional speed of moving, but more focused on the safety and standardization of the service process, as well as the protection of personal privacy and property safety. In order to meet these increasingly diverse needs, many transportation companies have increased their control over moving service personnel, striving to comprehensively improve the overall level of moving services through a series of management measures.

[0074] In the current moving service management and control system, the main focus is on the following aspects: First, in terms of process, it is necessary to strictly supervise the service behavior of movers, including but not limited to checking whether there are violations such as "improper packing", "maliciously guiding customers to modify packages", "guiding customers to cancel orders", "obtaining customer contact information" and "asking about customer privacy"; secondly, in terms of service, focus on the service attitude of movers to ensure that they do not have problems such as "negative service attitude" and "bad service attitude"; finally, on the platform red line, severe crackdowns will be taken on serious violations such as "private transactions" by movers.

[0075] However, the existing quality inspection methods for moving service violations have many pain points. Specifically, due to the huge amount of data generated during the moving service process, including pictures, text, voice and other types of information, the traditional manual quality inspection method is not only time-consuming and labor-intensive, but also inefficient, and it is difficult to ensure the accuracy of the judgment results. This manual quality inspection model is often unable to cope with large-scale data and cannot meet the needs of efficient and accurate quality inspection, thus affecting the improvement of the overall quality of moving services and the improvement of customer satisfaction.

[0076] Therefore, in view of the various problems existing in the existing quality inspection methods for violations of moving services, it is urgent to develop a more efficient and accurate quality inspection technology to achieve real-time monitoring and precise judgment of the behavior of moving service personnel, thereby comprehensively improving the quality of moving services and customer satisfaction.

[0077] Continue to refer Figure 2 , shows a flow chart of an embodiment of the data judgment and processing method proposed in this application. The embodiment of this application can acquire and process relevant data based on artificial intelligence technology.

[0078] The data judgment processing method provided by the embodiment of the present invention includes the following steps:

[0079] S210, obtaining multimodal data of movers in the moving service; wherein the multimodal data at least includes item packaging pictures, customer complaint texts, order evaluation texts, application voice information, and GPS track information.

[0080] In this step, the present invention can be applied to the quality inspection business scenarios of moving services, such as moving business service management, driver management, customer service quality inspection, order risk warning and other scenarios. The execution subject of the present invention can be specifically a judgment and processing system, which can be referred to as the system. The above-mentioned multimodal data refers to the multimodal information involved in the moving service process by the movers, including various types of text, pictures, voice, numerical information, etc. Specifically, (1) by deploying a camera device (such as a mobile phone, camera or special camera) at the moving site. During the moving process, the movers are required to take photos of the packaging of the items (whether the movers pack in a standardized manner during the moving process, whether the items are piled up or damaged), and upload them to the system. The system automatically stores and classifies these pictures for subsequent analysis. (2) By setting up customer complaint channels (such as telephone, email, complaint function in APP, etc.). Customers can submit complaint text information about the moving service through these channels (whether the customer has complaint information about this moving order). The system automatically records the complaint content, and performs preliminary classification and storage. (3) After the moving service is completed, invite the customer to evaluate the service. The customer provides evaluation text (whether the customer has any evaluation information for this moving order), including positive and negative feedback. The system collects and stores these evaluation texts for subsequent analysis. (4) During the moving process, use the APP's recording function to record the communication voice between the movers and the customer (the communication voice between the movers and the customer during the moving service). Ensure that the recording quality is clear for subsequent analysis. The system automatically stores these voice messages and classifies and marks them. (5) Install a GPS positioning device on the moving vehicle. Real-time collection of the moving vehicle's driving trajectory data, including time, longitude, latitude and other information, for example, including a series of numerical features, such as: vehicle ID, time, longitude, and latitude. Through these, it can be determined whether the mover walks according to the communicated time and established route during the service process, and whether there is a risk of lateness or delay. The system stores and analyzes the GPS data to determine whether the movers provide services according to the agreed time and route.

[0081] The specific implementation process of obtaining the multimodal data of movers in the moving service will be described in further detail in subsequent specific embodiments of the present invention, and will not be elaborated on here.

[0082] The present invention proposes a large-scale model-based moving service behavior quality inspection scheme. Through two-stage processing, a refined behavior quality inspection process is formed. The first stage is the extraction stage, which uses the self-trained multimodal large model to extract and generate key information in the multimodal information in a vertical manner, reducing redundant information and highlighting key features. The second stage is the summary stage, which uses the excellent summary generation ability of the quantitative large model and the core features extracted in the first stage to replace the idea of ​​existing technical classification, and obtain the quality inspection result of the moving service behavior, that is, the judgment result.

[0083] S220: Perform feature extraction on the multimodal data based on a preset feature extraction strategy to obtain corresponding multimodal features.

[0084] In this step, the specific implementation process of extracting features from the multimodal data based on a preset feature extraction strategy to obtain corresponding multimodal features will be further described in detail in subsequent specific embodiments of the present invention and will not be elaborated on here.

[0085] S230, performing feature extraction processing on the multimodal features based on a preset multimodal large model to obtain corresponding valid information.

[0086] In this step, the above-mentioned specific implementation process of performing feature extraction processing on the multimodal features based on the preset multimodal large model to obtain the corresponding effective information will be further described in detail in the subsequent specific embodiments of the present invention, and will not be elaborated in detail here.

[0087] S240, performing accountability analysis and processing on the effective information based on a preset quantitative large model to obtain accountability results corresponding to the movers.

[0088] In this step, the specific implementation process of performing responsibility analysis and processing on the effective information based on the preset quantitative large model to obtain the responsibility result corresponding to the mover will be further described in detail in the subsequent specific embodiments of the present invention and will not be elaborated here.

[0089] S250, outputting the judgment result.

[0090] In this step, the specific implementation process of training and evaluating the preset deep learning model based on the behavioral image sample data to obtain a behavior recognition model that meets the preset evaluation requirements will be further described in detail in the subsequent specific embodiments of the present invention and will not be elaborated on here.

[0091] In this step, the generated judgment results can be output in a suitable manner, such as generating a report, sending an email, etc. Feedback on the judgment results can also be collected for subsequent model optimization and improvement.

[0092] In an embodiment of the present invention, firstly, multimodal data of movers in moving services are obtained; wherein, the multimodal data at least includes pictures of item packaging, customer complaint texts, order evaluation texts, application voice information and GPS track information; then, feature extraction is performed on the multimodal data based on a preset feature extraction strategy to obtain corresponding multimodal features; then, feature refinement processing is performed on the multimodal features based on a preset multimodal big model to obtain corresponding valid information; subsequently, responsibility analysis processing is performed on the valid information based on a preset quantitative big model to obtain a responsibility result corresponding to the movers; finally, the responsibility result is outputted. The present invention obtains multimodal data of movers in moving services, then extracts features from the multimodal data based on a preset feature extraction strategy to obtain corresponding multimodal features, and then refines the multimodal features based on a multimodal big model to obtain corresponding effective information, and subsequently performs accountability analysis on the effective information based on a preset quantitative big model to obtain accountability results corresponding to the movers, and outputs the accountability results, thereby achieving comprehensive, accurate and efficient accountability processing for the moving service behaviors of the movers, improving the accountability processing efficiency for the moving service behaviors, and ensuring the accuracy of the obtained accountability results.

[0093] Optionally, the extracting features from the multimodal data based on a preset feature extraction strategy to obtain corresponding multimodal features includes:

[0094] Image feature extraction and text feature extraction are performed on the item packaging pictures respectively to obtain corresponding matrix features and text vector features.

[0095] In this step, for the item packaging pictures, the item packaging pictures will be converted into matrix features and attached with text vector features of the picture description. Among them, the item packaging pictures can be input into the convolutional neural network for feature extraction by loading a pre-trained convolutional neural network (such as VGG, ResNet, etc.), and the specified features can be extracted from the output layer of the convolutional neural network. The specified features can be matrix features (such as edges, textures, etc.) and text vector features (such as item names, categories, etc.).

[0096] Perform text feature extraction on the customer complaint text to obtain corresponding text features.

[0097] In this step, the customer complaint text will be converted into text features based on the text itself. Among them, text features such as keywords, sentiment tendencies, semantic relationships, etc. can be extracted from the above customer complaint text by using natural language processing (NLP) technology.

[0098] Emotional features are extracted from the order evaluation text to obtain a corresponding first emotional feature.

[0099] In this step, for the order evaluation text, a first emotional feature is extracted from the order evaluation text. The emotional feature of the order evaluation text can be extracted by using a sentiment analysis model, and the obtained first emotional feature can be used to analyze key issues and the degree of verbal conflict in the order evaluation.

[0100] Speech feature extraction and emotion feature extraction are performed on the application speech information to obtain corresponding speech features and second emotion features.

[0101] In this step, for the application voice information, the audio, timbre, waveform and other voice features are extracted from the application voice information, and converted into text through ASR technology, and then the service emotions are comprehensively evaluated from five fine-grained levels of "positive", "stable", "complaint", "anger" and "malicious attack", and the corresponding second emotional features are obtained. Among them, the application voice information can be converted into text by loading a pre-trained speech recognition model (such as ASR), and the audio, timbre, waveform and other voice features of the voice can be extracted using an audio processing library (such as Librosa), and then the text and voice features are combined, and the corresponding emotional features are extracted using a sentiment analysis model, thereby obtaining the second emotional features.

[0102] Numerical features are extracted from the GPS trajectory information to obtain corresponding numerical features.

[0103] In this step, the GPS trajectory information is statistically analyzed by using its numerical features to obtain the travel trajectory and risk characteristics. Among them, by statistically analyzing the GPS data, the travel trajectory, speed, stay time and other characteristics are extracted. And based on the extracted characteristics, it is judged whether the movers provide services according to the agreed time and route, and whether there are risky behaviors, so as to obtain the corresponding risk characteristics.

[0104] The matrix features, the text vector features, the text features, the first emotion features, the voice features, the second emotion features and the numerical features are integrated to obtain the multimodal features.

[0105] In this step, the matrix features, the text vector features, the text features, the first emotion features, the voice features, the second emotion features and the numerical features are integrated to obtain corresponding integrated features, and the integrated features are used as the above-mentioned multimodal features.

[0106] In an embodiment of the present invention, image feature extraction and text feature extraction are performed on the packaging picture of the items respectively to obtain corresponding matrix features and text vector features; and text feature extraction is performed on the customer complaint text to obtain corresponding text features; and emotion feature extraction is performed on the order evaluation text to obtain corresponding first emotion features; and voice feature extraction and emotion feature extraction are performed on the application voice information to obtain corresponding voice features and second emotion features; and numerical feature extraction is performed on the GPS track information to obtain corresponding numerical features; and subsequently the matrix features, the text vector features, the text features, the first emotion features, the voice features, the second emotion features and the numerical features are integrated to obtain the multimodal features. The present invention extracts features from the packaging pictures of items, customer complaint texts, order evaluation texts, application voice information and GPS track information contained in the multimodal data respectively, and then integrates all the extracted features, so that the corresponding multimodal features can be automatically and accurately obtained, ensuring the diversity and accuracy of the obtained multimodal features.

[0107] Optionally, the performing feature extraction processing on the multimodal features based on the preset multimodal macro model to obtain corresponding valid information includes:

[0108] The multimodal features are fused to obtain corresponding multimodal feature vectors.

[0109] In this step, the specific implementation process of fusing the multimodal features to obtain the corresponding multimodal feature vector will be further described in detail in subsequent specific embodiments of the present invention and will not be elaborated on here.

[0110] The multimodal feature vector is input into a characterization capability layer in the multimodal large model, and the multimodal feature vector is characterized by the characterization capability layer to obtain a corresponding first feature.

[0111] In this step, the multimodal large model is a pre-trained vertical large model, which includes at least a representation capability layer, an attention mechanism, and a decoder. The representation capability layer is a deep neural network that can perform high-level abstraction and representation of input features. By using the representation capability layer to perform high-level abstraction and representation processing on the input multimodal feature vector, the corresponding first feature is output.

[0112] The first feature is input into the attention application layer in the multimodal large model, and the first feature is subjected to feature selection processing by the attention application layer to obtain a corresponding second feature.

[0113] In this step, the attention application layer is a trainable neural network layer based on the attention mechanism, and the application of the attention mechanism can be used to evaluate the importance of features of different modalities. The attention application layer can assign a weight to each feature, indicating the importance of the feature in the current task. Among them, the above-mentioned input of the first feature into the attention application layer in the multimodal large model, the feature selection processing of the first feature by the attention application layer, and the specific implementation process of obtaining the corresponding second feature will be further described in detail in the subsequent specific embodiments of the present invention, and will not be elaborated on here.

[0114] The second feature is decoded based on the decoder in the multimodal large model to obtain a corresponding output sequence.

[0115] In this step, the second feature can be decoded and generated by using the decoder part of the multimodal large model to obtain the corresponding output sequence. The decoder can be a structure such as a recurrent neural network (RNN), a long short-term memory network (LSTM) or a Transformer, which can convert the feature vector into text or structured information that can be understood by humans. In addition, during the decoding process, strategies such as beam search can be used to find the most likely output sequence to ensure that the generated valid information is both accurate and complete.

[0116] The output sequence is used as the effective information.

[0117] In this step, in multimodal analysis, the features extracted by the feature fusion layer are used to generate effective information on different modalities based on these core features through decoding.

[0118] Image analysis: By analyzing the image data, the "regularity" and "damage" of the items during the handling and packaging process can be evaluated. The analysis of a large amount of image data can reveal operational errors or irregularities in the handling process. Analyzing the regularity of the items during the handling process helps to evaluate quality inspection.

[0119] Customer complaint analysis: Extract and analyze the "key issues" and "degree of verbal conflict" in customer complaints. By refining the core issues of customer complaints, we can accurately understand the customer's dissatisfaction with the service. These issues may involve details in the service process, the behavior of the movers, and the charges. Accurately identifying the root cause of the problem will help the model analyze the problem more accurately and make judgments.

[0120] Emotional analysis: Extract and analyze the "emotional state" between customers and movers. Through emotional analysis, we can gain insight into the emotional state of customers and movers in different service links, such as happiness, complaints, anxiety, dissatisfaction, etc., to help determine the service attitude of movers.

[0121] Extraction of text and voice violation information: Extract and analyze the "violations" involved in all text and voice modal information. By extracting and analyzing the call recordings between movers and customers, we can understand whether the movers' language when communicating with customers is polite and professional, and promptly discover and correct inappropriate language or service attitude to improve service quality. Extracting violations from voice information, such as service promises that are inconsistent with reality and opaque charging information, can effectively extract important violation information.

[0122] GPS violation information extraction: Extract and analyze "violations" in all GPS data. By analyzing GPS data, the actual driving route of the vehicle during the moving service can be monitored to detect violations such as deviation and unreasonable route. Analyze delays caused by illegal operations such as route deviation and speeding. Determine whether there are violations of service standards such as detours to increase fees and private parking.

[0123] The present invention can well extract the important points in the multimodal data by using a multimodal large model, that is, obtain the effective information corresponding to each mode for the next stage of summary and judgment.

[0124] In an embodiment of the present invention, the multimodal features are fused to obtain a corresponding multimodal feature vector; then the multimodal feature vector is input into a representation capability layer in the multimodal large model, and the multimodal feature vector is represented by the representation capability layer to obtain a corresponding first feature; then the first feature is input into an attention application layer in the multimodal large model, and the first feature is feature selected by the attention application layer to obtain a corresponding second feature; subsequently, the second feature is decoded based on a decoder in the multimodal large model to obtain a corresponding output sequence; finally, the output sequence is used as the effective information. The present invention obtains a corresponding multimodal feature vector by fusing multimodal features, then characterizes the multimodal feature vector based on the use of a characterization capability layer in a multimodal large model to obtain a corresponding first feature, and then performs feature selection on the first feature based on the use of an attention application layer in the multimodal large model to obtain a corresponding second feature, and subsequently decodes the second feature based on the use of a decoder in the multimodal large model, thereby achieving efficient and accurate feature extraction of multimodal features, improving the processing efficiency and accuracy of feature extraction, and ensuring the accuracy of the obtained effective information.

[0125] Optionally, the fusing the multimodal features to obtain a corresponding multimodal feature vector includes:

[0126] Get the preset fusion strategy.

[0127] In this step, there is no specific limitation on the selection of the above fusion strategy, which can be determined according to actual business needs. For example, any one of the strategies such as simple feature concatenation, weighted average or more complex fusion methods (such as tensor fusion) can be adopted.

[0128] The multimodal features are fused based on the fusion strategy to obtain corresponding fusion features.

[0129] In this step, the above multimodal features can be fused according to the strategy content of the selected fusion strategy to obtain fused features as the final multimodal feature vector.

[0130] The fused features are used as the multimodal feature vector.

[0131] In an embodiment of the present invention, a preset fusion strategy is obtained; then the multimodal features are fused based on the fusion strategy to obtain corresponding fusion features; and the fusion features are subsequently used as the multimodal feature vector. The present invention obtains a preset fusion strategy, and then fuses the multimodal features based on the use of the fusion strategy, so that the corresponding multimodal feature vector can be automatically and accurately generated, thereby ensuring the data accuracy of the obtained multimodal feature vector.

[0132] Optionally, the step of inputting the first feature into an attention application layer in the multimodal large model, performing feature selection processing on the first feature through the attention application layer to obtain a corresponding second feature includes:

[0133] The first feature is input into an attention application layer in the multimodal large model.

[0134] In this step, the attention application layer is a trainable neural network layer based on the attention mechanism. The attention application layer can assign a weight to each feature to indicate the importance of the feature in the current task.

[0135] Based on the attention application layer, a weight allocation process is performed on the first feature to obtain an output importance weight.

[0136] In this step, the attention application layer evaluates the importance of the first feature of the input and assigns a weight to each first feature, where the weight represents the importance of the feature in the current task.

[0137] The first features are screened and listed based on the importance weights to obtain corresponding initial features.

[0138] In this step, the above-mentioned first features can be screened and listed according to the output importance weights. Features with higher weights are considered to be more important and should be given priority for subsequent processing. Specifically, the first designated features whose importance weights are greater than the preset weight threshold can be screened out from the first features, and then the first designated features can be sorted in descending order of importance weights to obtain the sorted second designated features, and the second designated features can be used as the above-mentioned initial features. Among them, there is no specific limitation on the value of the above-mentioned weight threshold, which can be set according to actual business needs.

[0139] The initial features are combined to obtain a combined feature set.

[0140] In this step, the filtered initial features can be combined to form a new feature set, which contains the most important and relevant information in all modalities.

[0141] The feature set is used as the second feature.

[0142] In this step, the present invention does not use rough fusion by splicing, which will not bring negative gain to subsequent judgment. For example, simple splicing causes the features to be too long, and the front features often bring the risk of catastrophic forgetting, and the information in the previous sequence is forgotten during the model learning process. This application adopts the feature fusion properties of a multimodal large model, uses the attention mechanism to count the importance of each modal feature and the degree of connection between the modal features, screens and fuses them into multimodal features, and greatly refines the effective features for combination to avoid catastrophic forgetting.

[0143] In an embodiment of the present invention, the first feature is input into the attention application layer in the multimodal large model; then the first feature is weighted based on the attention application layer to obtain the output importance weight; then the first feature is screened and listed based on the importance weight to obtain the corresponding initial feature; the initial features are subsequently combined to obtain a combined feature set; and finally the feature set is used as the second feature. The present invention inputs the first feature into the attention application layer in the multimodal large model, then the first feature is weighted based on the attention application layer to obtain the output importance weight, and the first feature is screened and listed based on the importance weight to obtain the corresponding initial feature, and then the initial features are combined, so that the feature selection process for the first feature can be automatically and accurately completed, effectively ensuring that the generated second feature is the most important and relevant feature information in the first feature, and ensuring the accuracy of the generated second feature.

[0144] Optionally, the effective information is analyzed and processed based on a preset quantitative large model to obtain a judgment result corresponding to the mover, including:

[0145] The effective information is analyzed based on the quantitative big model to extract key information related to the judgment.

[0146] In this step, the effective information extracted by features can be input into the above-mentioned quantitative large model, so that the effective information can be deeply learned and analyzed by the quantitative large model to extract key information related to the judgment. Among them, the construction process of the above-mentioned quantitative large model may include: first, a deep learning model that has been quantified is selected, which has strong data processing and feature learning capabilities. Then, the model is trained using historical data and known judgment results so that it can learn the correlation between features and judgment results. Subsequently, the model is verified through a validation set to ensure that its performance and accuracy meet the requirements, thereby obtaining a trained quantitative large model.

[0147] The key information is converted into corresponding prompt words.

[0148] In this step, the key information can be converted into concise and clear prompt words by using the algorithms and mechanisms within the above-mentioned quantitative large model. The generated prompt words can accurately reflect the core points of the judgment results. These prompt words are important conclusions drawn by the quantitative large model after understanding and analyzing the input features, and they provide an important basis for subsequent judgments.

[0149] The prompt words are sorted and processed to obtain corresponding target prompt words.

[0150] In this step, the generated prompt words can be sorted and organized to ensure that they are logically clear and well-organized, so as to obtain the sorted target prompt words.

[0151] Call the preset judgment rules.

[0152] In this step, the above-mentioned accountability rules may include business rules and accountability standards. Among them, business rules may include the company's service standards, customer rights protection policies, etc., which provide clear guidance and basis for accountability. The accountability standards may involve specific division of responsibilities, compensation rules, etc., which are used to determine specific matters such as the attribution of responsibilities and the amount of compensation.

[0153] Based on the judgment rule, the target prompt word is subjected to judgment analysis and processing to obtain a judgment result corresponding to the mover.

[0154] In this step, after obtaining the target prompt word, a comprehensive judgment and analysis will be conducted in combination with the above-mentioned judgment rules. Specifically, it may include further interpretation of the target prompt word, analysis of the applicability of the business rules, and weighing of the judgment criteria. Finally, based on the results of the comprehensive judgment, a clear and specific judgment result is generated, and necessary explanations and instructions are included. Among them, the judgment result can be output in an appropriate manner, such as generating a report, sending a notification, etc., so that relevant personnel can understand and handle it. Specifically, the judgment result includes multiple dimensions, which can cover the entire scene content of the move, including "whether the packaging is standardized", "whether the customer is maliciously guided to modify the package", "whether the customer is guided to cancel", "whether the customer's contact information is obtained", "whether the customer's privacy is consulted", "whether the service attitude is negative", "whether the service attitude is bad", and "whether it is a private transaction". And each dimension has a clear judgment result (such as "yes" or "no"), and can be accompanied by relevant evidence or instructions.

[0155] In an embodiment of the present invention, by performing information analysis on the effective information based on the quantitative large model, key information related to the judgment is extracted; then the key information is converted into corresponding prompt words; then the prompt words are sorted out to obtain corresponding target prompt words; subsequently, preset judgment rules are called; finally, based on the judgment rules, the target prompt words are analyzed and processed for judgment, and the judgment results corresponding to the movers are obtained. The present invention performs information analysis on the effective information based on the use of the quantitative large model, extracts key information related to the judgment, then the key information is converted into corresponding prompt words, and the prompt words are sorted out to obtain corresponding target prompt words, and then based on the use of the judgment rules, the target prompt words are analyzed and processed for judgment, so that the judgment results corresponding to the movers can be obtained quickly and accurately, effectively improving the processing efficiency of the judgment analysis, and improving the accuracy of the generated judgment results.

[0156] Optionally, the obtaining of multimodal data of movers in the moving service includes:

[0157] Initial multimodal data of the movers in the moving service is obtained.

[0158] In this step, the initial multimodal data generated by the above-mentioned movers during the moving service without any processing can be collected, including but not limited to initial item packaging pictures, initial customer complaint texts, initial order evaluation texts, initial application voice information, initial GPS track information, etc.

[0159] A preprocessing strategy corresponding to the initial multimodal data is obtained.

[0160] In this step, the preprocessing strategy includes a first preprocessing strategy corresponding to the image data included in the initial multimodal data, and a second preprocessing strategy including the text data and voice data included in the initial multimodal data. The strategy content of the first preprocessing strategy may include: scaling, cropping, denoising and other processing to improve image quality. The strategy content of the second preprocessing strategy may include: removing irrelevant characters, word segmentation, noise reduction and other processing to improve analysis accuracy.

[0161] The initial multimodal data is preprocessed based on the preprocessing strategy to obtain corresponding designated multimodal data.

[0162] In this step, all different types of data included in the initial multimodal data may be preprocessed accordingly according to the policy content of the preprocessing policy, so as to obtain the preprocessed designated multimodal data as the final multimodal data.

[0163] The designated multimodal data is used as the multimodal data.

[0164] In an embodiment of the present invention, the initial multimodal data of the movers in the moving service is obtained; then a preprocessing strategy corresponding to the initial multimodal data is obtained; then the initial multimodal data is preprocessed based on the preprocessing strategy to obtain corresponding designated multimodal data; and the designated multimodal data is subsequently used as the multimodal data. The present invention obtains the initial multimodal data of the movers in the moving service, then a preprocessing strategy corresponding to the initial multimodal data is obtained, and then the initial multimodal data is preprocessed based on the use of the preprocessing strategy, so that the data quality and data accuracy of the generated multimodal data can be effectively ensured, which is conducive to improving the accuracy of subsequent analysis of the multimodal data.

[0165] In some optional implementations, the user information obtained is subject to the user's consent and complies with relevant laws and policies.

[0166] The present invention enhances the multimodal information fusion method: it not only fuses multimodal information in the feature space, but also enhances the deep collaboration and interaction between the modalities, thereby improving the feature fusion effect and the accuracy of judgment. At the same time, a more advanced fusion method is adopted to effectively capture the high-order interaction relationship between the modalities and enhance the accuracy and comprehensiveness of the quality inspection results.

[0167] Through refined feature processing: Before fusion, in-depth analysis is performed on the features of each modality to extract the core information and the correlation information between the modalities, rather than directly feeding the fused features into the classifier model. The importance information of each modality is identified and distinguished to avoid indiscriminate processing, so as to perform feature fusion more targetedly.

[0168] Through the innovative judgment process structure: adopting a two-stage approach, first refining the core features and then summarizing the final judgment results. Each stage is processed by using the excellent understanding and generation capabilities of the big model.

[0169] The present invention achieves the effects of improving the robustness of accountability, reducing costs and increasing efficiency, improving compliance, and reducing transaction violations, as follows:

[0170] 1. Improved robustness of judgment capability: The system's ability to process multimodal information in complex environments has been improved, making its application more extensive and stable. By enhancing the interactivity of modal information quality inspection, improving feature fusion effects, and capturing multimodal high-order feature relationships, the accuracy of quality inspection results is comprehensively improved.

[0171] 2. Reduce costs and increase efficiency: Reduce the cost of manual quality inspection and avoid the complaint costs caused by movers' complaints after quality inspection.

[0172] 3. Improve compliance: Through comprehensive model monitoring, it is easier for movers to be aware of violations being inspected, reduce the occurrence of violations, and enhance the awareness of behavioral norms among movers, thereby improving the overall violation rate and ensuring the company's service quality.

[0173] 4. Reduce transaction violations: Through a strict model detection mechanism, reduce the incidence of private transactions, improve the overall order completion rate, and increase the platform's order volume and revenue.

[0174] Further references Figure 3 , as a response to the above Figure 2 In order to realize the method shown in the figure, the present application provides an embodiment of a data judgment processing device 300, which is similar to Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0175] An embodiment of the present invention provides a data judgment processing device 300, and the data judgment processing device 300 includes:

[0176] The acquisition module 310 is used to acquire multimodal data of movers in the moving service; wherein the multimodal data at least includes pictures of item packaging, customer complaint text, order evaluation text, application voice information and GPS track information;

[0177] An extraction module 320, configured to extract features from the multimodal data based on a preset feature extraction strategy to obtain corresponding multimodal features;

[0178] A first processing module 330 is used to perform feature extraction processing on the multimodal features based on a preset multimodal large model to obtain corresponding valid information;

[0179] The second processing module 340 is used to perform accountability analysis on the effective information based on a preset quantitative large model to obtain an accountability result corresponding to the mover;

[0180] The output module 350 is used to output the judgment result.

[0181] Optionally, the extraction module 320 includes:

[0182] The first extraction submodule is used to extract image features and text features from the item packaging pictures to obtain corresponding matrix features and text vector features;

[0183] The second extraction submodule is used to extract text features from the customer complaint text to obtain corresponding text features;

[0184] A third extraction submodule is used to extract emotional features from the order evaluation text to obtain a corresponding first emotional feature;

[0185] A fourth extraction submodule, used for performing voice feature extraction and emotion feature extraction on the application voice information to obtain corresponding voice features and second emotion features;

[0186] A fifth extraction submodule is used to extract numerical features from the GPS trajectory information to obtain corresponding numerical features;

[0187] An integration submodule is used to integrate the matrix features, the text vector features, the text features, the first emotion features, the voice features, the second emotion features and the numerical features to obtain the multimodal features.

[0188] Optionally, the first processing module 330 includes:

[0189] A first processing submodule, configured to perform fusion processing on the multimodal features to obtain corresponding multimodal feature vectors;

[0190] A second processing submodule, configured to input the multimodal feature vector into a characterization capability layer in the multimodal large model, and perform characterization processing on the multimodal feature vector through the characterization capability layer to obtain a corresponding first feature;

[0191] A third processing submodule, configured to input the first feature into an attention application layer in the multimodal large model, and perform feature selection processing on the first feature through the attention application layer to obtain a corresponding second feature;

[0192] A fourth processing submodule, configured to decode the second feature based on a decoder in the multimodal large model to obtain a corresponding output sequence;

[0193] The first determining submodule is configured to use the output sequence as the valid information.

[0194] Optionally, the first processing submodule includes:

[0195] An acquisition unit, used for acquiring a preset fusion strategy;

[0196] A fusion unit, used for fusing the multimodal features based on the fusion strategy to obtain corresponding fusion features;

[0197] The first determining unit is configured to use the fused feature as the multimodal feature vector.

[0198] Optionally, the third processing submodule includes:

[0199] An input unit, used for inputting the first feature into an attention application layer in the multimodal large model;

[0200] an allocating unit, configured to perform a weight allocation process on the first feature based on the attention application layer to obtain an output importance weight;

[0201] A processing unit, configured to screen and list the first features based on the importance weights to obtain corresponding initial features;

[0202] A combining unit, used for combining the initial features to obtain a combined feature set;

[0203] The second determining unit is configured to use the feature set as the second feature.

[0204] Optionally, the second processing module 340 includes:

[0205] An analysis submodule, used to analyze the effective information based on the quantitative large model and extract key information related to the judgment;

[0206] A conversion submodule, used to convert the key information into corresponding prompt words;

[0207] A combing submodule, used for combing the prompt words to obtain corresponding target prompt words;

[0208] The calling submodule is used to call the preset judgment rules;

[0209] A generating submodule is used to perform responsibility analysis on the target prompt word based on the responsibility judgment rule to obtain a responsibility judgment result corresponding to the moving personnel.

[0210] Optionally, the acquisition module 310 includes:

[0211] A first acquisition submodule is used to acquire initial multimodal data of the mover in the moving service;

[0212] A second acquisition submodule, used to acquire a preprocessing strategy corresponding to the initial multimodal data;

[0213] A preprocessing submodule, configured to preprocess the initial multimodal data based on the preprocessing strategy to obtain corresponding designated multimodal data;

[0214] The second determining submodule is configured to use the designated multimodal data as the multimodal data.

[0215] To solve the above technical problems, the present application also provides a computer device. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.

[0216] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected through a system bus. It should be noted that the figure only shows a computer device 4 with components 41-43, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (Application Specific Integrated Circuit, ASIC), programmable gate arrays (Field-Programmable Gate Array, FPGA), digital processors (Digital Signal Processor, DSP), embedded devices, etc.

[0217] The computer device may be a computing device such as a desktop computer, a notebook, a PDA, a cloud server, etc. The computer device may interact with a user through a keyboard, a mouse, a remote controller, a touch pad, or a voice control device.

[0218] The memory 41 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as a hard disk or memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (FlashCard), etc. equipped on the computer device 4. Of course, the memory 41 can also include both the internal storage unit of the computer device 4 and its external storage device. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as the program code of the data judgment processing method, etc. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or are to be output.

[0219] The processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips in some embodiments. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to run the program code stored in the memory 41 or process data, such as running the program code of the data judgment processing method.

[0220] The network interface 43 may include a wireless network interface or a wired network interface. The network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.

[0221] The present application also provides another implementation, namely, providing a computer-readable storage medium, wherein the computer-readable storage medium stores the application crash processing program, and the application crash processing program can be executed by at least one processor to enable the at least one processor to perform the steps of the data judgment processing method as described above.

[0222] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware online platform, and of course, by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0223] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0224] Obviously, the embodiments described above are only some embodiments of the present application, rather than all embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application is described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions recorded in the aforementioned specific implementation methods, or to perform equivalent replacement of some of the technical features therein. Any equivalent structure made using the contents of the specification and drawings of this application, directly or indirectly used in other related technical fields, is similarly within the scope of patent protection of this application.

Claims

1. A data accountability processing method, characterized in that: include: Acquire multimodal data of movers in the moving service; wherein the multimodal data at least includes pictures of item packaging, customer complaint text, order evaluation text, application voice information and GPS track information; Extracting features from the multimodal data based on a preset feature extraction strategy to obtain corresponding multimodal features; Based on a preset multimodal large model, feature extraction processing is performed on the multimodal features to obtain corresponding effective information; Based on a preset quantitative large model, the effective information is analyzed and processed to obtain a judgment result corresponding to the moving personnel; The judgment result is outputted.

2. The method according to claim 1, characterized in that The extracting features of the multimodal data based on a preset feature extraction strategy to obtain corresponding multimodal features includes: Performing image feature extraction and text feature extraction on the item packaging pictures respectively to obtain corresponding matrix features and text vector features; Performing text feature extraction on the customer complaint text to obtain corresponding text features; Extracting emotional features from the order evaluation text to obtain a corresponding first emotional feature; Performing voice feature extraction and emotion feature extraction on the application voice information to obtain corresponding voice features and second emotion features; Extracting numerical features from the GPS trajectory information to obtain corresponding numerical features; The matrix features, the text vector features, the text features, the first emotion features, the voice features, the second emotion features and the numerical features are integrated to obtain the multimodal features.

3. The method according to claim 1, characterized in that The multimodal features are subjected to feature extraction processing based on the preset multimodal macro model to obtain corresponding valid information, including: Performing fusion processing on the multimodal features to obtain corresponding multimodal feature vectors; Inputting the multimodal feature vector into a characterization capability layer in the multimodal large model, performing characterization processing on the multimodal feature vector through the characterization capability layer to obtain a corresponding first feature; Inputting the first feature into the attention application layer in the multimodal large model, performing feature selection processing on the first feature through the attention application layer to obtain a corresponding second feature; Decoding the second feature based on the decoder in the multimodal large model to obtain a corresponding output sequence; The output sequence is used as the effective information.

4. The method according to claim 3, characterized in that The fusing process of the multimodal features to obtain a corresponding multimodal feature vector includes: Get the preset fusion strategy; Based on the fusion strategy, the multimodal features are fused to obtain corresponding fusion features; The fused features are used as the multimodal feature vector.

5. The method according to claim 3, characterized in that: The step of inputting the first feature into the attention application layer in the multimodal large model, performing feature selection processing on the first feature through the attention application layer to obtain a corresponding second feature includes: Inputting the first feature into an attention application layer in the multimodal large model; Performing weight allocation processing on the first feature based on the attention application layer to obtain an output importance weight; Screening and listing the first features based on the importance weights to obtain corresponding initial features; Combining the initial features to obtain a combined feature set; The feature set is used as the second feature.

6. The method according to claim 1, characterized in that The effective information is analyzed and processed based on the preset quantitative large model to obtain the accountability result corresponding to the mover, including: Performing information analysis on the effective information based on the quantitative big model to extract key information related to the judgment; Convert the key information into corresponding prompt words; Sorting out the prompt words to obtain corresponding target prompt words; Invoke the preset judgment rules; Based on the judgment rule, the target prompt word is subjected to judgment analysis and processing to obtain a judgment result corresponding to the mover.

7. The method according to claim 1, characterized in that The obtaining of multimodal data of movers in the moving service includes: Acquiring initial multimodal data of the mover in the moving service; Acquire a preprocessing strategy corresponding to the initial multimodal data; Preprocessing the initial multimodal data based on the preprocessing strategy to obtain corresponding designated multimodal data; The designated multimodal data is used as the multimodal data.

8. A data judgment processing device, characterized in that: include: An acquisition module is used to acquire multimodal data of movers in the moving service; wherein the multimodal data at least includes pictures of item packaging, customer complaint text, order evaluation text, application voice information and GPS track information; An extraction module, used to extract features from the multimodal data based on a preset feature extraction strategy to obtain corresponding multimodal features; A first processing module is used to perform feature extraction processing on the multimodal features based on a preset multimodal large model to obtain corresponding valid information; The second processing module is used to perform accountability analysis on the effective information based on a preset quantitative large model to obtain an accountability result corresponding to the mover; The output module is used to output the judgment result.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the data judgment processing method as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the data judgment processing method according to any one of claims 1 to 7 are implemented.