Text processing method and device, electronic equipment, storage medium and program product

By extracting and analyzing semantic features and character statistical features in vehicle data, and using text classification models, the problem of low accuracy in vehicle data classification in the prior art is solved, and more efficient data analysis is achieved.

CN119988624APending Publication Date: 2025-05-13ZEBRED NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411998397.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The classification method of vehicle data in the prior art is relatively low, making it difficult to effectively utilize data during subsequent analysis.

Method used

By obtaining the pending text of the associated vehicle, extracting its semantic features and character statistical features, and using the text classification model to determine the text content type to improve classification accuracy.

Benefits of technology

It improves the classification accuracy of vehicle data and can more accurately determine the text content type, thereby improving the effectiveness of data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988624A_ABST
    Figure CN119988624A_ABST
Patent Text Reader

Abstract

The invention provides a text processing method and device, electronic equipment, a storage medium and a program product. The method comprises the steps of obtaining a to-be-processed text of an associated vehicle; extracting semantic features and character statistical features of the to-be-processed text; and determining a text content type included in the to-be-processed text based on the semantic feature and the character statistical feature of the to-be-processed text. Through the method, the accuracy of the determined text content type can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of text processing, and in particular to a text processing method, device, electronic device, storage medium and program product. Background Art

[0002] With the development of information technology, the automotive industry has accumulated a huge amount of unstructured data sets, including vehicle maintenance records, user reviews, and driving data, which are crucial to improving vehicle quality and optimizing user experience. However, there are many types of vehicle-related data, and the data collection methods and storage locations are also diverse. Therefore, it is necessary to classify the vehicle data and obtain the type of each data so that targeted analysis can be performed based on the data type in subsequent analysis. However, the current classification method for vehicle data is less accurate. Summary of the invention

[0003] In order to overcome the problems existing in the related art, the present disclosure provides a text processing method, device, electronic device, storage medium and program product.

[0004] According to a first aspect of an embodiment of the present disclosure, a text processing method is provided, the method comprising:

[0005] Get the pending text of the associated vehicle;

[0006] Extracting semantic features and character statistical features of the text to be processed;

[0007] Based on the semantic features and character statistical features of the text to be processed, the text content type included in the text to be processed is determined.

[0008] In some embodiments, the method further comprises:

[0009] Segmenting the text to be processed to obtain multiple text segments;

[0010] The step of extracting the semantic features and character statistical features of the text to be processed includes:

[0011] For each text segment, extract the semantic features and character statistical features of the text segment;

[0012] The determining of the text content type included in the text to be processed based on the semantic features and character statistical features of the text to be processed includes:

[0013] For each text segment, based on the semantic features and character statistical features of the text segment, determine the text content type included in the text segment;

[0014] Based on the text content type included in each text segment, the text content type included in the to-be-processed text is determined.

[0015] In some embodiments, determining the text content type included in the text to be processed based on the semantic features and character statistical features of the text to be processed includes:

[0016] Based on the semantic features and character statistical features of the text to be processed, using a text classification model to determine the text content type included in the text to be processed;

[0017] The training method of the text classification model includes:

[0018] Acquire a sample data set; wherein the sample data set includes a plurality of sample text data, each sample text data being associated with a text type label;

[0019] For each sample text data, the semantic features and character statistical features of the sample text data are extracted, and the text type of the sample text data is predicted based on the semantic features and character statistical features of the sample text data using a preset neural network model;

[0020] Based on the difference between the text type predicted for each sample text data and the text type label of the sample text data, the neural network model is trained to obtain the text classification model.

[0021] In some embodiments, the plurality of sample text data are data including a plurality of preset characters; the method further comprises:

[0022] After the text classification model is obtained through training, the classification accuracy of sample text data associated with each preset character is counted based on the text classification model;

[0023] Determine the target character by selecting the preset characters that meet the preset accuracy condition among the classification accuracies associated with the preset characters;

[0024] The step of extracting the semantic features and character statistical features of the text to be processed includes:

[0025] The semantic features of the text to be processed and the statistical features of the target characters are extracted.

[0026] In some embodiments, the method further comprises:

[0027] Acquire a verification data set; wherein the verification data set includes a plurality of verification text data, each verification text data is associated with a text type label; the verification text data in the verification data set is different from the sample text data in the sample data set;

[0028] The method of training the neural network model based on the difference between the text type predicted for each sample text data and the text type label of the sample text data to obtain the text classification model includes:

[0029] Based on the difference between the text type predicted for each sample text data and the text type label of the sample text data, the neural network model is trained to obtain an initially trained neural network model;

[0030] For each verification text data, extracting semantic features and character statistical features of the verification text data, and using the initial neural network model to predict the text type of the verification text data based on the semantic features and character statistical features of the verification text data;

[0031] Based on the difference between the text type predicted by each verification text data and the text type label of the verification text data, the initially trained neural network model is fine-tuned to obtain the text classification model.

[0032] In some embodiments, the text to be processed includes at least one of the following:

[0033] Vehicle operating data;

[0034] User data of the user associated with the vehicle;

[0035] Sensor data associated with the vehicle;

[0036] Resource data associated with the vehicle.

[0037] In some embodiments, extracting the semantic features of the text to be processed includes:

[0038] The semantic features of the text to be processed are extracted using a bidirectional transformer encoding model.

[0039] According to a second aspect of an embodiment of the present disclosure, a text processing device is provided, the device comprising:

[0040] A first acquisition module is configured to acquire the to-be-processed text associated with the vehicle;

[0041] An extraction module, configured to extract semantic features and character statistical features of the text to be processed;

[0042] The first determination module is configured to determine the text content type included in the text to be processed based on the semantic features and character statistical features of the text to be processed.

[0043] In some embodiments, the text processing apparatus further comprises:

[0044] A segmentation module, configured to segment the text to be processed to obtain multiple text segments;

[0045] The extraction module is further configured to extract semantic features and character statistical features of each text segment;

[0046] The first determination module is further configured to determine, for each text segment, the text content type included in the text segment based on the semantic features and character statistical features of the text segment; and determine the text content type included in the text to be processed based on the text content type included in each text segment.

[0047] In some embodiments, the first determination module is further configured to determine the text content type included in the text to be processed using a text classification model based on the semantic features and character statistical features of the text to be processed;

[0048] Wherein, the training device of the text classification model includes:

[0049] A second acquisition module is configured to acquire a sample data set; wherein the sample data set includes a plurality of sample text data, each sample text data being associated with a text type label;

[0050] The prediction module is configured to extract the semantic features and character statistical features of each sample text data, and predict the text type of the sample text data based on the semantic features and character statistical features of the sample text data using a preset neural network model;

[0051] The training module is configured to train the neural network model based on the difference between the text type predicted for each sample text data and the text type label of the sample text data to obtain the text classification model.

[0052] In some embodiments, the plurality of sample text data are data including a plurality of preset characters; the text processing device further includes:

[0053] A statistical module, configured to, after training the text classification model, count the classification accuracy of sample text data associated with each preset character based on the text classification model;

[0054] A second determination module is configured to determine the target character by selecting the preset characters that meet the preset accuracy condition among the classification accuracies associated with the preset characters;

[0055] The extraction module is further configured to extract semantic features of the text to be processed and statistical features of the target characters.

[0056] In some embodiments, the training device of the text classification model further includes:

[0057] A third acquisition module is configured to acquire a verification data set; wherein the verification data set includes a plurality of verification text data, each verification text data is associated with a text type label; the verification text data in the verification data set is different from the sample text data in the sample data set;

[0058] The training module is further configured to train the neural network model based on the difference between the text type predicted for each sample text data and the text type label of the sample text data to obtain an initially trained neural network model; for each verification text data, extract the semantic features and character statistical features of the verification text data, and use the initial neural network model to predict the text type of the verification text data based on the semantic features and character statistical features of the verification text data; based on the difference between the text type predicted for each verification text data and the text type label of the verification text data, fine-tune the initially trained neural network model to obtain the text classification model.

[0059] In some embodiments, the text to be processed includes at least one of the following:

[0060] Vehicle operating data;

[0061] User data of the user associated with the vehicle;

[0062] Sensor data associated with the vehicle;

[0063] Resource data associated with the vehicle.

[0064] In some embodiments, the extraction module is further configured to extract semantic features of the text to be processed using a bidirectional transformer encoding model.

[0065] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, including:

[0066] processor;

[0067] Memory for storing computer programs or instructions;

[0068] The processor executes the computer program or instructions to implement the steps of the text processing method described in the first aspect above.

[0069] According to a fourth aspect of an embodiment of the present disclosure, a non-temporary computer-readable storage medium is provided, wherein the storage medium stores a computer program or instructions. When the computer program or instructions in the storage medium are executed by a processor, the steps of the text processing method described in the first aspect above are implemented.

[0070] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program or instructions, which, when executed by a processor, implements the steps of the text processing method described in the first aspect above.

[0071] The technical solution provided by the embodiments of the present disclosure may have the following beneficial effects:

[0072] In an embodiment of the present disclosure, an electronic device obtains a text to be processed of an associated vehicle, extracts semantic features and character statistical features of the text to be processed, and determines the type of text content included in the text to be processed based on the semantic features and character statistical features of the text to be processed, wherein extracting the semantic features of the text to be processed can determine the meaning of the text, and extracting the character statistical features of the text can determine information such as the length of the text, vocabulary richness, and text readability, and determining the type of text content included in the text to be processed based on the semantic features and character statistical features of the text to be processed can improve the accuracy of the determined text content type.

[0073] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0075] Figure 1 is a flowchart of a text processing method provided by an embodiment of the present disclosure;

[0076] Figure 2 It is a flowchart of a text classification model training method provided by an embodiment of the present disclosure;

[0077] Figure 3 It is a schematic diagram of the principle of the text processing method provided by the embodiment of the present disclosure.

[0078] Figure 4 It is a block diagram of a text processing device provided by an embodiment of the present disclosure.

[0079] Figure 5 It is a structural block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0080] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices consistent with some aspects of the present disclosure as detailed in the appended claims.

[0081] Figure 1 is a flowchart of a text processing method provided by an embodiment of the present disclosure, such as Figure 1 As shown, the method includes:

[0082] S11, obtaining the to-be-processed text of the associated vehicle;

[0083] S12, extracting semantic features and character statistical features of the text to be processed;

[0084] S13: Determine the text content type included in the text to be processed based on the semantic features and character statistical features of the text to be processed.

[0085] The text processing method provided in the embodiments of the present disclosure may be executed by terminal devices such as user equipment (UE), mobile devices, user terminals, mobile phones, tablet computers, personal digital assistants (PDAs), handheld devices, computing devices, vehicle-mounted devices, wearable devices, etc.; it may also be cloud devices such as cloud servers. The embodiments of the present disclosure do not limit the execution subject. For ease of description, the embodiments of the present disclosure are described with electronic devices as the execution subject.

[0086] In step S11, the electronic device obtains the text to be processed of the associated vehicle, wherein the text to be processed of the associated vehicle is all data related to the vehicle. The electronic device can obtain the text to be processed by means of a web crawler; the electronic device can also call the text to be processed through a predetermined programming interface (Application Programming Interface, API); the electronic device can also obtain the text to be processed from a database or data set storing the text to be processed, such as a log data set of the vehicle, through a data processing tool; the electronic device can also establish a communication connection with the Internet of Vehicles platform to obtain the text to be processed from the Internet of Vehicles platform, and the embodiments of the present disclosure are not limited to this.

[0087] In some embodiments, the text to be processed includes at least one of the following:

[0088] Vehicle operating data;

[0089] User data of the user associated with the vehicle;

[0090] Sensor data associated with the vehicle;

[0091] Resource data associated with the vehicle.

[0092] In the disclosed embodiment, the text to be processed may include the vehicle's operating data, wherein the vehicle's operating data includes data generated by the vehicle during driving, such as vehicle speed, engine speed, fuel tank level, driving route, number of stops, etc. The vehicle's operating data is usually stored in the vehicle's log data set.

[0093] In the disclosed embodiment, the text to be processed may also include user data of vehicle-associated users. The vehicle-associated users may be the driver of the vehicle and / or the owner of the vehicle. The user data of the vehicle-associated users may include the driver's behavior data during the driving of the vehicle, such as the length of time the steering wheel is held, driving habit data, etc. The user data of the vehicle-associated users may also include interaction data between the user and the vehicle-mounted equipment, such as the interaction records between the user and the vehicle system (such as voice assistants, touch screens, etc.), such as voice commands, touch operations, etc., and the user's usage of vehicle-related applications (such as vehicle-mounted applications, remote control applications, etc.), such as frequency of use, dwell time, operating behavior, etc.; the user data of the vehicle-associated users may also include user information of the user, such as user name, gender, contact information, etc.; the user data of the vehicle-associated users may also include data uploaded by the user, such as vehicle maintenance records uploaded by the user, vehicle performance evaluation data, vehicle appearance evaluation data and other evaluation data.

[0094] In the disclosed embodiment, the text to be processed may also include vehicle-associated sensor data, wherein the vehicle-associated sensors may include perception sensors, such as cameras, radars, etc. The vehicle-associated sensors may also include positioning sensors, such as the Global Positioning System (GPS), Inertial Measurement Unit (IMU), etc. The vehicle-associated sensor data refers to data measured by the vehicle-associated sensors, such as environmental information during vehicle driving, vehicle location information, vehicle status information, etc.

[0095] In the disclosed embodiments, the text to be processed may also include resource data associated with the vehicle, wherein the resource data associated with the vehicle includes the vehicle's own hardware resource data, such as the model of the vehicle, data of various components of the vehicle, such as data of the vehicle's tires, engine, chassis, etc.; the resource data associated with the vehicle may also include the vehicle's software resource data, such as the vehicle's operating system's operating data, update data, the vehicle's associated application's operating data, update data, etc.; the resource data associated with the vehicle may also include the vehicle's network resource data, etc., such as the vehicle's network connection status data, the vehicle's network protocol, network IP address data, etc.

[0096] In some embodiments, the text to be processed may be plain English text data, such as vehicle operation data stored in a vehicle log data set. The sensor data associated with the vehicle is usually plain English text data, which may include English characters, numeric characters, and special symbol characters. Exemplarily, the text data to be processed may be "VehicleID:ABC123, Speed:60mph", wherein Vehicle ID refers to the identification information of the vehicle, which is the resource data associated with the vehicle, and Speed ​​refers to the speed of the vehicle, which is the operation data of the vehicle.

[0097] In other embodiments, the text to be processed may also be Chinese text data. For example, the data uploaded by the user may be Chinese text. For example, the text to be processed may be "This car has good acceleration performance and low fuel consumption, but the interior is a bit simple, and overall it is cost-effective." The text to be processed is the evaluation data uploaded by the user on the vehicle performance and the interior decoration of the vehicle. In other embodiments, the text to be processed may also be Chinese and English text data including both English and Chinese characters, and the text to be processed may also be text data in other languages, which is not limited by the embodiments of the present disclosure.

[0098] In some embodiments, the electronic device can directly obtain the text to be processed, such as the operation data of the vehicle. The resource data associated with the vehicle is usually generated in the form of text data, which can be stored in the log data set of the vehicle. The electronic device can directly obtain the text to be processed from the log data set of the vehicle. In other embodiments, the data associated with the vehicle, such as the data format of the data uploaded by the user, may be image data, audio data, etc. At this time, the electronic device can obtain the image data or audio data, and convert the image data or audio data into the text to be processed for subsequent processing. In this regard, the embodiments of the present disclosure are not limited.

[0099] In the disclosed embodiment, after the electronic device obtains the text to be processed of the associated vehicle, it can first determine the language of the text to be processed. For example, when the text to be processed is Chinese text data, the text to be processed can be preprocessed by removing irrelevant characters such as punctuation marks and spaces, removing repeated phrases, removing noise in the text to be processed, such as HTML tags, JavaScript codes, etc., and converting the text to be processed into a preset format. When the text to be processed is English text data, the text to be processed can be preprocessed by removing unnecessary spaces, exemplarily replacing multiple consecutive spaces with one space, and switching all English characters in the English text to uppercase or lowercase mode.

[0100] In step S12, the electronic device extracts semantic features of the text to be processed. In some embodiments, the electronic device may generate a vector corresponding to the text to be processed using an embedding model, and the vector can represent the semantic features of the text to be processed.

[0101] In other embodiments, the electronic device may extract semantic features corresponding to the text to be processed based on the bag-of-words model, such as determining the number of times each phrase appears in the text to be processed, and analyzing the subject or keywords of the text to be processed based on the number of times each phrase appears in the text to be processed, thereby determining the semantic features of the text to be processed.

[0102] In some other embodiments, extracting the semantic features of the text to be processed includes:

[0103] The semantic features of the text to be processed are extracted using a bidirectional transformer encoding model.

[0104] In an embodiment of the present disclosure, an electronic device uses a bidirectional encoder representations from transformers (BERT) model to extract semantic features of a text to be processed, wherein the BERT model is a pre-trained language model based on a transformer architecture, and the BERT model uses a bidirectional transformer to capture contextual information of words. The BERT model divides the text to be processed into multiple phrases, and extracts and fuses the semantic features of each phrase to obtain the semantic features of the text to be processed, wherein the BERT model can simultaneously consider the text content before and after each phrase when extracting the semantic features of the text to be processed, thereby more accurately understanding the semantics of the phrase.

[0105] In the disclosed embodiment, the electronic device uses the BERT model to extract semantic features of the text to be processed. When extracting semantic features, the BERT model not only considers the semantic information of each phrase in the text to be processed, but also combines the contextual information of each phrase in the text to be processed, thereby improving the accuracy of the semantic features of the text to be processed.

[0106] In step S12, the electronic device also extracts character statistical features of the text to be processed, wherein the character statistical features may include the length of the text to be processed, and the length of the text to be processed refers to the total number of characters included in the text to be processed. For example, when the text to be processed is "Hello, world", the length of the text to be processed is 11.

[0107] In the embodiment of the present disclosure, the character statistical features may also include the frequency of each character in the text to be processed, wherein the frequency of each character may be determined by the number of occurrences of each character in the text to be processed, and the representation method may adopt a specific number of occurrences, or may adopt a description method such as high frequency, medium frequency, low frequency, etc., and the embodiment of the present disclosure does not impose any limitation on this. For example, when the text to be processed is "Hello, world", the character statistical features may be that the frequency of 'h' is 1, the frequency of 'e' is 1, the frequency of 'l' is 3, the frequency of 'o' is 2, the frequency of '' (space) is 1, the frequency of 'w' is 1, the frequency of 'r' is 1, and the frequency of 'd' is 1.

[0108] In the embodiments of the present disclosure, the character statistical feature may also include the frequency of preset characters in the text to be processed, wherein the preset characters are characters set in advance, such as numbers, @ characters, etc. For example, when the text to be processed is "VehicleID:ABC123, Speed:60mph", when the preset characters are numbers, the character statistical feature may be 5 (the text to be processed includes a total of 5 numeric characters). The frequency of the preset characters may also be expressed by the ratio of the preset characters to the total number of characters in the text to be processed. The embodiments of the present disclosure do not limit the representation form of the frequency of the preset characters.

[0109] In the disclosed embodiment, the character statistical features may also include the number of different characters in the text to be processed, the distribution of phrases of different lengths in the text to be processed, the compactness of characters in the text to be processed, etc.

[0110] In step S13, the electronic device determines the type of text content included in the text to be processed based on the semantic features and character statistical features of the text to be processed. In some embodiments, after the electronic device determines the semantic features and character statistical features of the text to be processed, it can perform weighted fusion on the semantic features and character statistical features to obtain weighted fused features, and determine the type of text content included in the text to be processed based on the weighted fused features, wherein the weights of the semantic features and character statistical features can be set values.

[0111] In other embodiments, the electronic device may first generate a first vector that can characterize the semantic features of the text to be processed, and extract the character statistical features of the text to be processed, merge the character statistical features of the text to be processed with the first vector to obtain a second vector that can simultaneously characterize the semantic features and character statistical features of the text to be processed, and determine the text content type included in the text to be processed based on the second vector.

[0112] In other embodiments, the electronic device may first generate a first vector that can characterize the semantic features of the text to be processed, extract character statistical features of the text to be processed, mark the character statistical features on the first vector in the form of labels, and determine the text content type included in the text to be processed based on the marked first vector.

[0113] In some embodiments, the electronic device may determine the type of text content included in the text to be processed based on the semantic features and character statistical features of the text to be processed using a preset classification model, wherein the preset classification model may be a support vector machine model (SVM), a logistic regression model, a decision tree model, a naive Bayes model, etc.

[0114] In other embodiments, the electronic device may determine the type of text content included in the text to be processed based on the semantic features and character statistical features of the text to be processed using a preset algorithm, wherein the preset algorithm may include a keyword matching method, a regular expression method, a naive Bayes method, and the like.

[0115] In an embodiment of the present disclosure, the electronic device determines the type of text content included in the text to be processed based on the semantic features and character statistical features of the text to be processed. For example, when the semantic features of the text to be processed represent that the text to be processed is user information, and the character statistical features of the text to be processed represent that the number of numeric characters in the text to be processed is 11, then it can be determined that the type of text content included in the text to be processed is the user contact information (telephone number) of a vehicle-associated user; when the semantic features of the text to be processed represent that the text to be processed is user information, and the character statistical features of the text to be processed represent that the number of @ characters in the text to be processed is 1, then it can be determined that the type of text content included in the text to be processed is the email address of the vehicle-associated user.

[0116] In an embodiment of the present disclosure, an electronic device obtains a text to be processed of an associated vehicle, extracts semantic features and character statistical features of the text to be processed, and determines the type of text content included in the text to be processed based on the semantic features and character statistical features of the text to be processed, wherein extracting the semantic features of the text to be processed can determine the meaning of the text, and extracting the character statistical features of the text can determine information such as the length of the text, vocabulary richness, and text readability, and determining the type of text content included in the text to be processed based on the semantic features and character statistical features of the text to be processed can improve the accuracy of the determined text content type.

[0117] In some embodiments, the method further comprises:

[0118] Segmenting the text to be processed to obtain multiple text segments;

[0119] The step of extracting the semantic features and character statistical features of the text to be processed includes:

[0120] For each text segment, extract the semantic features and character statistical features of the text segment;

[0121] The determining of the text content type included in the text to be processed based on the semantic features and character statistical features of the text to be processed includes:

[0122] For each text segment, based on the semantic features and character statistical features of the text segment, determine the text content type included in the text segment;

[0123] Based on the text content type included in each text segment, the text content type included in the to-be-processed text is determined.

[0124] In the embodiment of the present disclosure, the electronic device segments the text to be processed to obtain multiple text segments. Before the electronic device segments the text to be processed, it may first determine the language of the text to be processed. When the text to be processed is Chinese text, in some embodiments, the electronic device may segment the Chinese text based on the number of characters, and segment the Chinese text into multiple text segments. For example, when the Chinese text is "This car has good acceleration performance and low fuel consumption, but the interior is slightly shabby, and overall it is cost-effective.", the Chinese text may be segmented based on the number of characters so that the number of characters in each text segment is the same. For example, if the number of characters in each text segment is set to 6, the Chinese text may be segmented into "This car has good acceleration", "good performance, low fuel consumption", "but the interior", "slightly shabby", "overall cost-effective", and "high cost-effective." 6 text segments. It should be noted that each Chinese character and each punctuation mark in the Chinese text are regarded as a character, and the characters in each text segment of the Chinese text are continuous, and the characters in adjacent text segments may overlap or may not overlap. This is not limited by the embodiment of the present disclosure.

[0125] In other embodiments, the electronic device may also segment the Chinese text based on punctuation marks, and segment the Chinese text into multiple text segments, wherein when segmenting based on punctuation marks, the punctuation mark may be the last character in the text segment, or may be the first character in the text segment. For example, when the Chinese text is "This car has a very good acceleration performance and low fuel consumption, but the interior is a bit simple, and overall it is cost-effective.", the Chinese text may be segmented based on punctuation marks. Taking the punctuation mark as the last character in each text segment as an example, the Chinese text may be segmented into four text segments: "This car has a very good acceleration performance," "Low fuel consumption," "But the interior is a bit simple," and "Overall it is cost-effective." It should be noted that when segmenting the Chinese text based on punctuation marks, it is necessary to determine that the Chinese text includes punctuation marks. For example, when the Chinese text is pre-processed, the punctuation marks in the Chinese text are not removed.

[0126] In other embodiments, the electronic device may also segment the Chinese text based on special characters, where the special characters are set characters, such as " " (space) ".", etc., which is not limited in the embodiments of the present disclosure.

[0127] In the embodiments of the present disclosure, when the electronic device determines that the text to be processed is an English text, in some embodiments, the English text can be segmented based on punctuation marks, and the English text can be segmented into multiple text segments. For example, when the English text is "Vehicle ID: ABC123, Speed: 60mph", the English text can be segmented based on punctuation marks, and the English text can be segmented into four text segments of "Vehicle ID:", "ABC123," "Speed:", and "60mph". It should be noted that when segmenting based on punctuation marks, the punctuation mark can be the last character in the text segment or the first character in the text segment, and the embodiments of the present disclosure do not limit this.

[0128] In other embodiments, the electronic device may also segment the English text based on special characters, wherein the special characters are set characters, such as “ ” (space), “@”, etc., and the embodiments of the present disclosure do not limit this. For example, when the English text is “Vehicle ID: ABC123, Speed: 60mph”, and the special character is “:”, the English text may be segmented into three text segments: “Vehicle ID:”, “ABC123, Speed:”, and “60mph”. It should be noted that when segmenting based on special characters, the special character may be the last character in the text segment or the first character in the text segment, and the embodiments of the present disclosure do not limit this.

[0129] In the embodiments of the present disclosure, when the electronic device determines that the text to be processed is a Chinese-English text that includes both Chinese and English, in some embodiments, the electronic device can process the Chinese-English text, separate the Chinese and English in the Chinese-English text, and then segment the Chinese in the manner of the aforementioned Chinese text and segment the English in the manner of the aforementioned English text.

[0130] In other embodiments, the segmentation may be based on a segmentation method supported by both Chinese text and English text, such as segmentation based on punctuation marks or segmentation based on special characters, and the embodiments of the present disclosure do not limit this.

[0131] In the disclosed embodiment, the electronic device extracts the semantic features and character statistical features of each text segment, wherein, as mentioned above, the electronic device may generate a vector corresponding to each text segment using an embedding model, and the vector may represent the semantic features of each text segment. The electronic device may also extract the semantic features corresponding to each text segment based on a bag-of-words model. The electronic device may also extract the semantic features of each text segment based on a BERT model.

[0132] In the disclosed embodiment, the character statistical features of each text fragment may include the features described above, such as the length of each text fragment, the frequency of each character in each text fragment, the frequency of preset characters in each text fragment, the number of different characters in each text fragment, the distribution of phrases of different lengths in each text fragment, the compactness of characters in each text fragment, etc.

[0133] In the disclosed embodiments, the electronic device determines, for each text segment, the type of text content included in the text segment based on the semantic features and character statistical features of the text segment. As mentioned above, for each text segment, the electronic device can determine, based on the semantic features and character statistical features of the text segment, the type of text content included in the text segment using a preset classification model; the electronic device can also determine the type of text content included in the text segment using a preset algorithm.

[0134] In an embodiment of the present disclosure, the electronic device determines the text content type included in the text to be processed based on the text content type included in each text fragment. In some embodiments, the electronic device may simply merge the text content type included in each text fragment based on the text content type included in each text fragment to obtain the text content type included in the text to be processed. Exemplarily, the text to be processed includes two text fragments, the text content type included in text fragment 1 is the contact information of a vehicle-associated user, and the text content type included in text fragment 2 is the vehicle model. Then the text content type included in the text to be processed is the contact information of the vehicle-associated user and the model of the vehicle.

[0135] In the disclosed embodiment, the electronic device segments the text to be processed to obtain multiple text segments, extracts semantic features and character statistical features of each text segment, determines the text content type included in each text segment based on the semantic features and character statistical features of each text segment, and determines the text content type included in the text to be processed based on the text content type included in each text segment. On the one hand, the segmented text segments contain more specific and local text information, which can more accurately reflect the text content type included in the text segments when analyzed separately. For example, the text to be processed may have multiple text content types, and a global analysis of the text to be processed may be interfered by context information, resulting in classification errors. Therefore, segmenting the text to be processed can more accurately determine the content type of each text segment, thereby further improving the accuracy of the determined text content type of the text to be processed. On the other hand, the text to be processed may contain sensitive characters, such as the user name of the associated vehicle. These sensitive characters may not be so conspicuous or easily ignored in the text to be processed. By segmenting the text to be processed and focusing on each text segment, the electronic device can more accurately identify these sensitive characters.

[0136] In some embodiments, determining the text content type included in the text to be processed based on the semantic features and character statistical features of the text to be processed includes:

[0137] Based on the semantic features and character statistical features of the text to be processed, using a text classification model to determine the text content type included in the text to be processed;

[0138] in, Figure 2 is a flowchart of a text classification model training method provided by an embodiment of the present disclosure, such as Figure 2 As shown, the training method of the text classification model includes:

[0139] S21, obtaining a sample data set; wherein the sample data set includes a plurality of sample text data, each sample text data being associated with a text type label;

[0140] S22, for each sample text data, extracting semantic features and character statistical features of the sample text data, and using a preset neural network model to predict the text type of the sample text data based on the semantic features and character statistical features of the sample text data;

[0141] S23. Based on the difference between the text type predicted for each sample text data and the text type label of the sample text data, the neural network model is trained to obtain the text classification model.

[0142] In the disclosed embodiment, the electronic device determines the type of text content included in the text to be processed based on the semantic features and character statistical features of the text to be processed and uses a text classification model. As mentioned above, after the electronic device determines the semantic features and character statistical features of the text to be processed, it can perform weighted fusion of the semantic features and the character statistical features to obtain weighted fused features, and input the weighted fused features into the text classification model to output the type of text content included in the text to be processed; the electronic device can also fuse the character statistical features of the text to be processed with a first vector representing the semantic features of the text to be processed to obtain a second vector that can simultaneously represent the semantic features and the character statistical features of the text to be processed, and input the second vector into the text classification model to output the type of text content included in the text to be processed.

[0143] In step S21, the electronic device obtains a sample data set, wherein the sample data set includes a plurality of sample text data, each of which is associated with a text type label. In some embodiments, the electronic device may obtain historical text data of the associated vehicle, and use the historical text data as sample text data, and each of which is associated with a text type label, wherein the text type label associated with the historical text data may be manually determined, or may be determined using a preset classification model or classification method, and the embodiments of the present disclosure do not limit this.

[0144] In other embodiments, the electronic device may obtain text data associated with the vehicle input by the user, and a text type label corresponding to the text data input by the user. In other embodiments, the electronic device may also obtain a sample data set based on a generation model such as Generative Adversarial Networks (GANs), such as GANs may generate text data associated with the vehicle and a text type label corresponding to each text data, wherein the text type label is used to indicate the type of text content included in the sample text data.

[0145] In the disclosed embodiment, after the electronic device obtains the sample data set, it can clean the sample data set, such as removing noise data and outliers, processing missing data, etc., and can also convert each sample data in the sample data set into the same format.

[0146] In step S22, the electronic device extracts the semantic features and character statistical features of each sample text data, wherein, as mentioned above, the electronic device can use an embedding model to generate a vector corresponding to each sample text data, and the vector can represent the semantic features of the sample text data. The electronic device can also extract the semantic features corresponding to each sample text data based on a bag-of-words model. The electronic device can also extract the semantic features of each sample text data based on a BERT model.

[0147] In the embodiments of the present disclosure, the character statistical features of each sample text data may include the features described above, such as the length of each sample text data, the frequency of each character in each sample text data, the frequency of preset characters in each sample text data, the number of different characters in each sample text data, the distribution of phrases of different lengths in each sample text data, the compactness of characters in each sample text data, etc.

[0148] In the disclosed embodiment, the electronic device extracts the semantic features and character statistical features of the sample text data for each sample text data, and then uses a preset neural network model to predict the text type of the sample text data based on the semantic features and character statistical features of the sample text data. After the electronic device extracts the semantic features and character statistical features of the sample text data, it may first pre-process the extracted semantic features and character statistical features of the sample text data so that the neural network model can understand and perform subsequent processing, the pre-processing including removing irrelevant characters, converting to corresponding vector representations, etc. The neural network model predicts the text type of the sample text data based on the semantic features and character statistical features of the pre-processed sample text data.

[0149] In step S23, the electronic device trains the neural network model based on the difference between the text type predicted for each sample text data and the text type label of the sample text data to obtain a text classification model, wherein training the neural network model means adjusting the parameters of the neural network model to obtain the model parameters of the trained neural network model as the model parameters of the text classification model, wherein the electronic device can determine the difference between the text type predicted for each sample text data and the text type label of the sample text data through methods such as bilingual evaluation understudy (BLEU), cosine similarity, and Euclidean distance, and use the difference as a loss value to adjust the parameters of the neural network model, such as using an optimization algorithm to update the parameters of the neural network model, and the optimization algorithm can be a gradient descent method, an adaptive learning rate (Adaptive Moment Estimation, Adam) method, and the like.

[0150] In the disclosed embodiment, the text content type included in the text to be processed is determined using a trained text classification model, which can further improve the accuracy of the determined text content type.

[0151] In some embodiments, the plurality of sample text data are data including a plurality of preset characters;

[0152] The method further comprises:

[0153] After the text classification model is obtained through training, the classification accuracy of sample text data associated with each preset character is counted based on the text classification model;

[0154] Determine the target character by selecting the preset characters that meet the preset accuracy condition among the classification accuracies associated with the preset characters;

[0155] The step of extracting the semantic features and character statistical features of the text to be processed includes:

[0156] The semantic features of the text to be processed and the statistical features of the target characters are extracted.

[0157] In the embodiments of the present disclosure, the multiple sample text data are data including multiple preset characters. In some embodiments, the multiple sample text data include sample text data with different text contents or the multiple sample text data include sample text data with different character types. For example, the multiple sample text data include sample text data including numbers and sample text data including the @ character.

[0158] In other embodiments, the sample data set includes a group of sample text data, and different preset characters are set. The character statistical features of this group of sample text data for different preset characters are counted to obtain multiple groups of sample text data including multiple preset characters. For example, the character statistical features of this group of sample text data for the @ character can be counted to obtain a first sample text data group, and the character statistical features of this group of sample text data for the * character can be counted to obtain a second sample text data group. The first sample text data group and the second sample text data group constitute multiple sample text data. It should be noted that the number of preset characters each time is not necessarily unique. For example, the character statistical features of this group of sample text data for both the @ character and the * character can be counted to obtain a third sample text data group. The embodiments of the present disclosure do not limit this.

[0159] In other embodiments, the sample data set may also include multiple groups of sample text data, and different preset characters are set. The character statistical features of the multiple groups of sample text data for different preset characters are counted to obtain multiple groups of sample text data including multiple preset characters. The specific statistical method is as above.

[0160] In the disclosed embodiment, when extracting character statistical features of sample text data, the electronic device may continuously change the representation form of the character statistical features to improve the flexibility of the text classification model.

[0161] In an embodiment of the present disclosure, after training a text classification model, the electronic device counts the classification accuracy of sample text data associated with each preset character based on the text classification model, wherein the classification accuracy of sample text data associated with each group of preset characters based on the text classification model can be determined based on the ratio of sample text data in which the difference between the text type predicted by the text classification model and the text type label of the sample text data is less than a preset difference threshold and all sample text data associated with each group of preset characters.

[0162] In the disclosed embodiment, after the electronic device determines the classification accuracy of the sample text data associated with each preset character based on the text classification model, the preset character that meets the preset accuracy condition among the classification accuracy associated with each preset character is determined as the target character. Among them, the preset character that meets the preset accuracy condition among the classification accuracy associated with each preset character can be the preset character with the highest classification accuracy.

[0163] In the disclosed embodiment, when the electronic device extracts the semantic features and character statistical features of the text to be processed, it extracts the semantic features of the text to be processed and the statistical features of the target character. For example, when the target character is the @ character, the electronic device extracts the character statistical features of the text to be processed, which means that the electronic device determines the frequency of occurrence of the @ character in the text to be processed.

[0164] In the disclosed embodiment, the electronic device selects a preset character whose accuracy meets the preset accuracy condition as the target character, and when extracting the character statistical features of the text to be processed, extracts the statistical features of the target character in the text to be processed. Since the target character is a preset character whose accuracy meets the preset accuracy condition, therefore, extracting the statistical features of the target character in the text to be processed can, on the one hand, improve the accuracy of the model during actual use, and on the other hand, there is no need to extract the statistical features of each character in the text to be processed, which can speed up the extraction of the statistical features of the characters of the text to be processed and is highly intelligent.

[0165] In some embodiments, the method further comprises:

[0166] Acquire a verification data set; wherein the verification data set includes a plurality of verification text data, each verification text data is associated with a text type label; the verification text data in the verification data set is different from the sample text data in the sample data set;

[0167] The method of training the neural network model based on the difference between the text type predicted for each sample text data and the text type label of the sample text data to obtain the text classification model includes:

[0168] Based on the difference between the text type predicted for each sample text data and the text type label of the sample text data, the neural network model is trained to obtain an initially trained neural network model;

[0169] For each verification text data, extracting semantic features and character statistical features of the verification text data, and using the initial neural network model to predict the text type of the verification text data based on the semantic features and character statistical features of the verification text data;

[0170] Based on the difference between the text type predicted by each verification text data and the text type label of the verification text data, the initially trained neural network model is fine-tuned to obtain the text classification model.

[0171] In an embodiment of the present disclosure, an electronic device obtains a verification data set; wherein the verification data set includes multiple verification text data, each verification text data is associated with a text type label, and the verification text data in the verification data set is different from the sample text data in the sample data set. In some embodiments, the electronic device may obtain historical text data of an associated vehicle, and use the historical text data as verification text data, and each historical text data is associated with a text type label, wherein the text type label associated with the historical text data may be manually determined, or may be determined using a preset classification model or classification method, and the embodiments of the present disclosure do not limit this. In other embodiments, the electronic device may obtain text data of an associated vehicle input by a user, and a text type label corresponding to the text data input by the user, and use the text data of an associated vehicle input by the user as verification text data.

[0172] In an embodiment of the present disclosure, the electronic device trains a neural network model based on the difference between the text type predicted for each sample text data and the text type label of the sample text data to obtain an initially trained neural network model, wherein the training method for training the neural network model using the sample text data is as described above.

[0173] In the disclosed embodiment, the electronic device extracts the semantic features and character statistical features of each verification text data, and uses the initially trained neural network model to predict the text type of the verification text data based on the semantic features and character statistical features of the verification text data. As mentioned above, the electronic device can use an embedded model to generate a vector corresponding to each verification text data, which can characterize the semantic features of the verification text data. The electronic device can also extract the semantic features corresponding to each verification text data based on a bag of words model. The electronic device can also extract the semantic features of each verification text data based on a BERT model.

[0174] In the embodiments of the present disclosure, the character statistical features of each verification text data may include the features described above, such as the length of each verification text data, the frequency of each character in each verification text data, the frequency of preset characters in each verification text data, the number of different characters in each verification text data, the distribution of phrases of different lengths in each verification text data, the closeness of characters in each verification text data, etc.

[0175] In the disclosed embodiment, the electronic device extracts the semantic features and character statistical features of the verification text data for each verification text data, and then uses the initially trained neural network model to predict the text type of the verification text data based on the semantic features and character statistical features of the verification text data. After the electronic device extracts the semantic features and character statistical features of the verification text data, it may first pre-process the extracted semantic features and character statistical features of the verification text data so that the initially trained neural network model can understand and perform subsequent processing, the pre-processing including removing irrelevant characters, converting to corresponding vector representations, etc. The initially trained neural network model predicts the text type of the verification text data based on the semantic features and character statistical features of the pre-processed verification text data.

[0176] In an embodiment of the present disclosure, the electronic device fine-tunes the initially trained neural network model based on the difference between the text type predicted by each verification text data and the text type label of the verification text data to obtain a text classification model, wherein the fine-tuning of the initially trained neural network model is to adjust the parameters of the initially trained neural network model to obtain the model parameters of the initially trained neural network model after adjustment, which are used as the model parameters of the text classification model, wherein the electronic device can determine the difference between the text type predicted by each verification text data and the text type label of the verification text data through methods such as the Bilingual Evaluation Understudy (BLEU) method, the cosine similarity method, and the Euclidean distance method, and use the difference as a loss value to adjust the parameters of the initially trained neural network model, such as using an optimization algorithm to update the parameters of the neural network model, and the optimization algorithm can be a gradient descent method, an adaptive learning rate (Adaptive Moment Estimation, Adam) method, and the like.

[0177] In the embodiment of the present disclosure, the electronic device may also determine the ratio of the number of verification text data in the verification data set in which the difference between the text type predicted by the verification text data and the text type label of the verification text data is less than a preset difference threshold to the number of all verification text data included in the verification data set, and adjust the parameters of the initially trained neural network model, wherein adjusting the parameters of the initially trained neural network model includes hyperparameter adjustment and adjustment of the structure of the initially trained neural network model.

[0178] In the disclosed embodiment, the electronic device evaluates the model trained based on the sample data set based on verification text data that is different from the sample text data in the sample data set, and adjusts the model parameters again based on the evaluation results of the verification data set, which can reduce the overfitting of the model to the sample data set and further improve the accuracy of the text classification model in determining the text content type.

[0179] Figure 3 : is a schematic diagram of the principle of the text processing method provided by the embodiment of the present disclosure, wherein text L31 is the text to be processed, L32 is the encoder layer in the BERT model, the semantic vector L33 represents the semantic features of the text L31, the text features L34, L35, and L36 are all character statistical features of the text L31, L37 is the structure in the fine-tuned text classification model, and L38 is the classification result of the text L31, representing the text content type included in the text L31. Figure 3 As shown, after the electronic device obtains the text L31, the text L31 is input into the encoder layer L32 in the BERT model to extract the semantic vector L33 that can represent the semantic features of the text L31, and extract the text features L34, L35, and L36 of the text L31, wherein each text feature may represent different features, such as the text feature L34 may represent the length of the text L31, the text feature L35 may represent the number of occurrences of numeric characters in the text L31, the text feature L36 may represent the number of occurrences of the @ character in the text L31, etc., and the semantic vector L33 and the text features L34, L35, and L36 are input into the text classification model L37, and the text classification model L37 determines the classification result L38 of the text L31 based on the semantic vector L33 and the text features L34, L35, and L36.

[0180] In the embodiments of the present disclosure, it should be noted that in the process of determining the text content type included in the text to be processed using the fine-tuned text classification model, users, such as developers of the text classification model, can regularly check the text content type included in the text to be processed determined by the text classification model, and determine the accuracy of the text content type included in the text to be processed determined by the text classification model. If the accuracy is less than a preset accuracy threshold, the text classification model is updated, wherein updating the text classification model refers to adjusting the model parameters of the text classification model based on the difference between the text content type included in the text to be processed determined by the text classification model and the text content type included in the text to be processed determined by the user.

[0181] Figure 4 4 is a block diagram of a text processing device 400 provided in an embodiment of the present disclosure. Figure 4 As shown, the device mainly includes:

[0182] A first acquisition module 401 is configured to acquire a to-be-processed text associated with a vehicle;

[0183] An extraction module 402 is configured to extract semantic features and character statistical features of the text to be processed;

[0184] The first determination module 403 is configured to determine the text content type included in the text to be processed based on the semantic features and character statistical features of the text to be processed.

[0185] In some embodiments, the text processing device 400 further includes:

[0186] A segmentation module, configured to segment the text to be processed to obtain multiple text segments;

[0187] The extraction module 402 is further configured to extract semantic features and character statistical features of each text segment;

[0188] The first determination module 403 is further configured to determine, for each text segment, the text content type included in the text segment based on the semantic features and character statistical features of the text segment; and determine the text content type included in the text to be processed based on the text content type included in each text segment.

[0189] In some embodiments, the first determination module 403 is further configured to determine the text content type included in the text to be processed using a text classification model based on the semantic features and character statistical features of the text to be processed;

[0190] Wherein, the training device of the text classification model includes:

[0191] A second acquisition module is configured to acquire a sample data set; wherein the sample data set includes a plurality of sample text data, each sample text data being associated with a text type label;

[0192] The prediction module is configured to extract the semantic features and character statistical features of each sample text data, and predict the text type of the sample text data based on the semantic features and character statistical features of the sample text data using a preset neural network model;

[0193] The training module is configured to train the neural network model based on the difference between the text type predicted for each sample text data and the text type label of the sample text data to obtain the text classification model.

[0194] In some embodiments, the plurality of sample text data are data including a plurality of preset characters; the text processing device 400 further includes:

[0195] A statistical module, configured to, after training the text classification model, count the classification accuracy of sample text data associated with each preset character based on the text classification model;

[0196] A second determination module is configured to determine the target character by selecting the preset characters that meet the preset accuracy condition among the classification accuracies associated with the preset characters;

[0197] The extraction module 402 is further configured to extract semantic features of the text to be processed and statistical features of the target characters.

[0198] In some embodiments, the training device of the text classification model further includes:

[0199] A third acquisition module is configured to acquire a verification data set; wherein the verification data set includes a plurality of verification text data, each verification text data is associated with a text type label; the verification text data in the verification data set is different from the sample text data in the sample data set;

[0200] The training module is further configured to train the neural network model based on the difference between the text type predicted for each sample text data and the text type label of the sample text data to obtain an initially trained neural network model; for each verification text data, extract the semantic features and character statistical features of the verification text data, and use the initial neural network model to predict the text type of the verification text data based on the semantic features and character statistical features of the verification text data; based on the difference between the text type predicted for each verification text data and the text type label of the verification text data, fine-tune the initially trained neural network model to obtain the text classification model.

[0201] In some embodiments, the text to be processed includes at least one of the following:

[0202] Vehicle operating data;

[0203] User data of the user associated with the vehicle;

[0204] Sensor data associated with the vehicle;

[0205] Resource data associated with the vehicle.

[0206] In some embodiments, the extraction module 402 is further configured to extract semantic features of the text to be processed using a bidirectional transformer encoding model.

[0207] Figure 5 is a structural block diagram of an electronic device 500 provided by an embodiment of the present disclosure. For example, the electronic device 500 may be a mobile phone, a computer, a digital broadcast terminal, a message transceiver device, a tablet device, a personal digital assistant, an electronic device in a vehicle, such as an in-vehicle communication system, a vehicle information management system, a vehicle driving assistance system, and the like.

[0208] Reference Figure 5, the electronic device 500 may include one or more of the following components: a processing component 502 , a memory 504 , a power component 506 , a multimedia component 508 , an audio component 510 , an input / output (I / O) interface 512 , a sensor component 514 , and a communication component 516 .

[0209] The processing component 502 generally controls the overall operation of the electronic device 500, such as operations associated with at least one of display, phone calls, data communications, camera operations, and recording operations. The processing component 502 may include one or more processors 520 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 502 may include one or more modules to facilitate the interaction between the processing component 502 and other components. For example, the processing component 502 may include a multimedia module to facilitate the interaction between the multimedia component 508 and the processing component 502.

[0210] The memory 504 is configured to store various types of data to support operations on the electronic device 500. Examples of such data include at least one of the following: instructions for any application or method operating on the electronic device 500, contact data, phone book data, messages, pictures, and videos. The memory 504 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk, or an optical disk.

[0211] The power supply component 506 provides power to various components of the electronic device 500. The power supply component 506 may include at least one of the following: a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 500.

[0212] The multimedia component 508 includes a screen that provides an output interface between the electronic device 500 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 508 includes a front camera and / or a rear camera. When the electronic device 500 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0213] The audio component 510 is configured to output and / or input audio signals. For example, the audio component 510 includes a microphone (MIC), and when the electronic device 500 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 504 or sent via the communication component 516. In some embodiments, the audio component 510 also includes a speaker for outputting audio signals.

[0214] I / O interface 512 provides an interface between processing component 502 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, a home button, a volume button, a start button, and a lock button.

[0215] The sensor assembly 514 includes one or more sensors for providing various aspects of status assessment for the electronic device 500. For example, the sensor assembly 514 can detect the open / closed state of the electronic device 500, the relative positioning of the components, such as the display and keypad of the electronic device 500, and the sensor assembly 514 can also detect the position change of the electronic device 500 or a component in the electronic device 500, the presence or absence of contact between the user and the electronic device 500, the orientation or acceleration / deceleration of the electronic device 500, and the temperature change of the electronic device 500. The sensor assembly 514 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 514 may also include a light sensor, such as a complementary metal oxide semiconductor (CMOS) or a charge coupled device (CCD) image sensor, for use in imaging applications. In some embodiments, the sensor assembly 514 may also include, but is not limited to, at least one of the following: an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, and a temperature sensor.

[0216] The communication component 516 is configured to facilitate communication between the electronic device 500 and other devices in a wired or wireless manner. The electronic device 500 can access a wireless network based on a communication standard, such as Wi-Fi, 4G, 5G, or a combination thereof. In an exemplary embodiment, the communication component 516 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 516 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0217] In an exemplary embodiment, the electronic device 500 can be implemented by one or more application specific integrated circuits (ASIC), digital signal processors (DSP), digital signal processing devices (DSPD), programmable logic devices (PLD), field programmable gate arrays (FPGA), controllers, microcontrollers, microprocessors or other electronic components.

[0218] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 504 including executable instructions or a computer program, which can be executed by a processor 520 of an electronic device 500 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.

[0219] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform any of the above-mentioned text processing methods of the embodiments of the present disclosure. For example, the method includes:

[0220] Get the pending text of the associated vehicle;

[0221] Extracting semantic features and character statistical features of the text to be processed;

[0222] Based on the semantic features and character statistical features of the text to be processed, the text content type included in the text to be processed is determined.

[0223] The embodiment of the present disclosure provides a computer program product, which includes: a computer program or executable instructions, which are stored in a computer-readable storage medium. The processor of the computer device reads the computer program or executable instructions from the computer-readable storage medium, and the processor executes the computer program or executable instructions, so that the computer device executes any one of the above-mentioned text processing methods of the embodiment of the present disclosure.

[0224] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0225] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A text processing method, characterized in that: The method comprises: Get the pending text of the associated vehicle; Extracting semantic features and character statistical features of the text to be processed; Based on the semantic features and character statistical features of the text to be processed, the text content type included in the text to be processed is determined.

2. The method according to claim 1, characterized in that The method further comprises: Segmenting the text to be processed to obtain multiple text segments; The step of extracting the semantic features and character statistical features of the text to be processed includes: For each text segment, extract the semantic features and character statistical features of the text segment; The determining of the text content type included in the text to be processed based on the semantic features and character statistical features of the text to be processed includes: For each text segment, based on the semantic features and character statistical features of the text segment, determine the text content type included in the text segment; Based on the text content type included in each text segment, the text content type included in the to-be-processed text is determined.

3. The method according to claim 1 or 2, characterized in that: The determining of the text content type included in the text to be processed based on the semantic features and character statistical features of the text to be processed includes: Based on the semantic features and character statistical features of the text to be processed, using a text classification model to determine the text content type included in the text to be processed; The training method of the text classification model includes: Acquire a sample data set; wherein the sample data set includes a plurality of sample text data, each sample text data being associated with a text type label; For each sample text data, the semantic features and character statistical features of the sample text data are extracted, and the text type of the sample text data is predicted based on the semantic features and character statistical features of the sample text data using a preset neural network model; Based on the difference between the text type predicted for each sample text data and the text type label of the sample text data, the neural network model is trained to obtain the text classification model.

4. The method according to claim 3, characterized in that The plurality of sample text data are data including a plurality of preset characters; The method further comprises: After the text classification model is obtained through training, the classification accuracy of sample text data associated with each preset character is counted based on the text classification model; Determine the target character by selecting the preset characters that meet the preset accuracy condition among the classification accuracies associated with the preset characters; The step of extracting the semantic features and character statistical features of the text to be processed includes: The semantic features of the text to be processed and the statistical features of the target characters are extracted.

5. The method according to claim 3, characterized in that: The method further comprises: Acquire a verification data set; wherein the verification data set includes a plurality of verification text data, each verification text data is associated with a text type label; the verification text data in the verification data set is different from the sample text data in the sample data set; The method of training the neural network model based on the difference between the text type predicted for each sample text data and the text type label of the sample text data to obtain the text classification model includes: Based on the difference between the text type predicted for each sample text data and the text type label of the sample text data, the neural network model is trained to obtain an initially trained neural network model; For each verification text data, extracting semantic features and character statistical features of the verification text data, and using the initial neural network model to predict the text type of the verification text data based on the semantic features and character statistical features of the verification text data; Based on the difference between the text type predicted by each verification text data and the text type label of the verification text data, the initially trained neural network model is fine-tuned to obtain the text classification model.

6. The method according to claim 1, characterized in that The text to be processed includes at least one of the following: Vehicle operating data; User data of the user associated with the vehicle; Sensor data associated with the vehicle; Resource data associated with the vehicle.

7. The method according to claim 1, characterized in that The extracting of semantic features of the text to be processed includes: The semantic features of the text to be processed are extracted using a bidirectional transformer encoding model.

8. A text processing device, characterized in that: The device comprises: A first acquisition module is configured to acquire the to-be-processed text associated with the vehicle; An extraction module, configured to extract semantic features and character statistical features of the text to be processed; The first determination module is configured to determine the text content type included in the text to be processed based on the semantic features and character statistical features of the text to be processed.

9. An electronic device, characterized in that: include: processor; Memory for storing computer programs or instructions; The processor executes the computer program or instructions to implement the steps of the text processing method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing a computer program or instruction, characterized in that: When the computer program or instructions in the storage medium are executed by a processor, the steps of the text processing method according to any one of claims 1 to 7 are implemented.

11. A computer program product, comprising a computer program or instructions, characterized in that: When the computer program or instruction is executed by a processor, the steps of the text processing method according to any one of claims 1 to 7 are implemented.