Text classification method and device, electronic equipment and storage medium

By introducing a target intention recognition model in text classification, extracting text features associated with spatiotemporal features, and determining intent based on specified spatiotemporal conditions, the shortcomings of intention recognition and classification in the prior art are solved, and higher classification accuracy and adaptability are achieved.

CN120030157APending Publication Date: 2025-05-23联想诺谛(北京)智能科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411897670.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify and classify different consent maps caused by changes in spatiotemporal information.

Method used

By obtaining the text data to be identified, the target text that represents the target text associated with the spatiotemporal features is extracted using the target intention recognition model, and the target intent is determined based on the target text and the specified spatiotemporal conditions. The model includes a text feature extraction module and a spatiotemporal feature extraction module. Through training scheme and loss function optimization, the parameters of the feature extraction module are adjusted to improve the accuracy of the model.

Benefits of technology

It improves the accuracy of text classification, can accurately identify and classify different consents under changes in space-time information, and enhances the adaptability and generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030157A_ABST
    Figure CN120030157A_ABST
Patent Text Reader

Abstract

The invention provides a text classification method and device, electronic equipment and a storage medium. The method comprises the steps of obtaining to-be-recognized text data; a target intention recognition model is used for recognizing the to-be-recognized text data, a target text in the to-be-recognized text data is extracted, and the target text represents a text associated with spatial-temporal characteristics; and based on the target text and the specified space-time condition, determining a target intention corresponding to the to-be-recognized text data under the specified space-time condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a text classification method, device, electronic device and storage medium. Background Art

[0002] As time and space information change, there may be different perspectives on the text. Therefore, how to determine the different intentions due to changes in time and space information has become a technical problem that needs to be solved urgently. Summary of the invention

[0003] The present disclosure provides a text classification method, device, electronic device and storage medium.

[0004] According to a first aspect of the present disclosure, a text classification method is provided, the method comprising:

[0005] Obtaining text data to be recognized;

[0006] Using a target intention recognition model to recognize the text data to be recognized, and extracting a target text from the text data to be recognized, wherein the target text is a text associated with a temporal and spatial feature;

[0007] Based on the target text and the specified time and space conditions, the target intent corresponding to the text data to be recognized under the specified time and space conditions is determined.

[0008] In an implementation of the present application, the training scheme of the target intent recognition model includes:

[0009] Acquire sample text data and label information of the sample text data, wherein the label information represents the real intention corresponding to the sample text data;

[0010] Inputting the sample text data into a model to be trained, wherein the model to be trained includes a text feature extraction module and a spatiotemporal feature extraction module;

[0011] The text feature extraction module extracts text features of the sample text data;

[0012] The spatiotemporal feature extraction module extracts the spatiotemporal features of the sample text data;

[0013] The model to be trained identifies the predicted intent corresponding to the sample text data based on the text features and the spatiotemporal features, and determines a first intent contribution rate corresponding to the text features and a second intent contribution rate corresponding to the spatiotemporal features;

[0014] Based on the label information and the predicted intent, determining a loss function value of the model to be trained;

[0015] Determining whether the model to be trained has completed training based on the loss function value, the first intention contribution rate, and the second intention contribution rate;

[0016] The model obtained after training is determined as the target intent recognition model.

[0017] In an implementation of the present application, determining whether the model to be trained has completed training based on the loss function value, the first intention contribution rate, and the second intention contribution rate includes:

[0018] If the loss function value is not less than a preset loss threshold, comparing the first intention contribution rate and the second intention contribution rate;

[0019] If the first intention contribution rate is greater than the second intention contribution rate, adjusting the parameters of the text feature extraction module of the model to be trained;

[0020] If the second intention contribution rate is greater than the first intention contribution rate, adjusting the parameters of the spatiotemporal feature extraction module of the model to be trained;

[0021] For new sample text data, return to the step of inputting the sample text data into the model to be trained until the loss function value is less than the preset loss threshold, and determine whether the model to be trained has completed training.

[0022] In an implementation of the present application, the determining of the first intention contribution rate corresponding to the text feature and the second intention contribution rate corresponding to the spatiotemporal feature includes:

[0023] Determine a first prediction classification result corresponding to the sample text data based on the text feature;

[0024] Determine a second prediction classification result corresponding to the sample text data according to the spatiotemporal feature;

[0025] Based on the prediction intention, determining weight information of the first prediction classification result and the second prediction classification result;

[0026] According to the weight of the first classification result and the weight of the second classification result, a first intention contribution rate corresponding to the text feature and a second intention contribution rate corresponding to the spatiotemporal feature are determined respectively.

[0027] In one implementation of the present application, the specified spatiotemporal conditions include at least one of time conditions, space conditions, and state conditions.

[0028] In an implementation of the present application, the method further includes:

[0029] The target intent recognition model is updated based on the target intent and the target text.

[0030] According to a second aspect of the present disclosure, a text classification device is provided, the device comprising:

[0031] A data acquisition module, used to acquire text data to be recognized;

[0032] A data recognition module, used to recognize the text data to be recognized by using a target intention recognition model, and extract a target text from the text data to be recognized, wherein the target text is a text associated with a temporal and spatial feature;

[0033] The intention determination module is used for determining the target intention corresponding to the text data to be identified under the specified time and space conditions based on the target text and the specified time and space conditions.

[0034] In an implementation of the present application, the device further includes:

[0035] A model training module is used to obtain sample text data and label information of the sample text data, wherein the label information represents the real intent corresponding to the sample text data; the sample text data is input into a model to be trained, wherein the model to be trained includes a text feature extraction module and a spatiotemporal feature extraction module; the text feature extraction module extracts the text features of the sample text data; the spatiotemporal feature extraction module extracts the spatiotemporal features of the sample text data; the model to be trained identifies the predicted intent corresponding to the sample text data based on the text features and the spatiotemporal features, and determines a first intent contribution rate corresponding to the text features and a second intent contribution rate corresponding to the spatiotemporal features; based on the label information and the predicted intent, determines a loss function value of the model to be trained; based on the loss function value, the first intent contribution rate and the second intent contribution rate, determines whether the model to be trained has completed training; and determines the model obtained after completing the training as a target intent recognition model.

[0036] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0037] at least one processor; and

[0038] a memory communicatively coupled to the at least one processor;

[0039] An image acquisition device; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the present disclosure.

[0040] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method described in the present disclosure.

[0041] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, in which:

[0043] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts.

[0044] Figure 1 A schematic diagram of an implementation process of the text classification method provided in an embodiment of the present application is shown;

[0045] Figure 2 A schematic diagram of a training process of a target intent recognition model provided in an embodiment of the present application is shown;

[0046] Figure 3 Another training process diagram of the target intent recognition model provided in the embodiment of the present application is shown;

[0047] Figure 4 A schematic diagram of the structure of a model to be trained provided in an embodiment of the present application is shown;

[0048] Figure 5 A schematic diagram of the structure of a text classification device provided in an embodiment of the present application is shown;

[0049] Figure 6 A schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0050] In order to make the purpose, features, and advantages of the present disclosure more obvious and easy to understand, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.

[0051] As time and space information change, there may be different perspectives on the text. Specifically, since the user's intentions and preferences in the classification scenario will change with the change of time and space information, in order to improve the accuracy of classification, the present application provides a text classification method, device, electronic device and storage medium for the situation where the user's intentions change dynamically with time and space information. The electronic device provided in the present application can be a mobile phone, a computer, a tablet computer and other devices.

[0052] The technical solution of the embodiment of the present application will be described below in conjunction with the drawings in the embodiment of the present application.

[0053] Figure 1 A schematic diagram of an implementation process of the text classification method provided in the embodiment of the present application is shown as follows: Figure 1 As shown, the method includes:

[0054] S101, obtaining text data to be recognized.

[0055] In the present disclosure, the text data to be identified may be the text data involved in the scenario where text classification is required. The text data to be identified may include document text data, web page text data, text data in pictures, and text data converted from speech, for example, text in emails, tweets on social media, and images.

[0056] S102, using a target intention recognition model to recognize the text data to be recognized, and extracting a target text from the text data to be recognized, wherein the target text is a text characterizing text associated with spatiotemporal features.

[0057] In the present disclosure, spatiotemporal features may include time features and / or space features. The time features of a text include time-related information in the text content. In one possible implementation, the time features of a text may include one or more of a timestamp, a time interval, a time series, timeliness information, a time tag, a periodic time, historical background information, and seasonal information. For example, the time feature of the text data "the best spots for cherry blossom viewing in spring" is the seasonal information "spring". For another example, the time feature of the text data "the badminton competition will be held from September 8, 2024 to September 10, 2024" is the competition time "September 8, 2024 to September 10, 2024".

[0058] The spatial features of the text include information related to the spatial location and region in the text content. In one possible implementation, the temporal features of the text may include one or more of the geographical location, spatial relationship, city information, spatial vocabulary, regional description information, etc. For example, the spatial feature in the text data "I want to go to Beijing" is the city information "Beijing" involved in the text. For another example, the spatial feature in the text data "G Province will experience a lunar eclipse in June 2023" is "G Province".

[0059] In the present disclosure, the target intent recognition model is obtained by training the to-be-trained model through sample text data and label information of the sample text data. The label information of the sample text data represents the real intent corresponding to the sample text data.

[0060] S103, based on the target text and the specified time and space conditions, determining the target intent corresponding to the text data to be recognized under the specified time and space conditions.

[0061] In the present disclosure, the specified spatiotemporal conditions may include at least one of time conditions, space conditions, and state conditions. The time conditions included in the specified spatiotemporal conditions may be information such as a preset time or a preferred time interval in the text classification scenario. For example, for a computer maintenance scenario, the user's expectations include working days, and the time condition of working days can be used as the specified spatiotemporal conditions. The spatial conditions included in the specified spatiotemporal conditions may be the area information or location information involved in the text classification scenario. For example, for a computer maintenance scenario, the maintenance frequency of users in the A department of Company B is the highest among all departments, and the spatial condition of the A department of Company B can be used as the specified spatiotemporal conditions.

[0062] The method provided by the embodiment of the present disclosure is adopted to obtain the text data to be identified, and the target intention recognition model is used to identify the text data to be identified, and the target text in the text data to be identified is extracted. The target text is a text that represents text associated with spatiotemporal features. Based on the target text and the specified spatiotemporal conditions, the target intention corresponding to the text data to be identified under the specified spatiotemporal conditions is determined. That is, in view of the situation that the user's intention changes dynamically with the spatiotemporal information, by extracting the target text related to the spatiotemporal features in the text data to be identified, the spatiotemporal features of the text data to be identified and the specified spatiotemporal conditions are used to accurately determine the target intention corresponding to the text data to be identified under the specified spatiotemporal conditions, thereby improving the accuracy of text classification.

[0063] In one possible implementation, Figure 2 A schematic diagram of a training process of a target intention recognition model provided in an embodiment of the present application is shown. Figure 2 As shown, the training scheme of the target intent recognition model includes:

[0064] S201, obtaining sample text data and label information of the sample text data, wherein the label information represents the real intention corresponding to the sample text data.

[0065] In the present disclosure, the label information of the sample text data represents the real intention corresponding to the sample text data. The label information of the sample text data may include user attitude information, information about the event that the user wants to do, and user expectation information. For example, if the sample text data is "I am very tired at work recently, and I want to go to Shanghai to relax in May", it can be determined that the label information of the sample text data is the user's real intention "to travel to Shanghai"; if the sample text data is "I have severe skin allergies after using product A, and I give a bad review", it can be determined that the label information of the sample text data is the user's real intention "negative attitude".

[0066] S202, inputting the sample text data into a model to be trained, wherein the model to be trained includes a text feature extraction module and a spatiotemporal feature extraction module.

[0067] In the present disclosure, the model to be trained may be a deep learning network, such as a convolutional neural network, a recurrent neural network, or a long short-term memory network.

[0068] In the present disclosure, the model to be trained includes a text feature extraction module and a spatiotemporal feature extraction module. The text feature extraction module may include a convolution layer, a pooling layer, and a fully connected layer. For example, for a text classification scenario, the convolution layer of the text feature extraction module may perform convolution processing on a text vector based on convolution kernels of different sizes to extract local features, the pooling layer of the text feature extraction module may extract the most significant features from the output features of the convolution layer, and the fully connected layer of the text feature extraction module may flatten the output features of the pooling layer and classify them through the fully connected layer.

[0069] The spatiotemporal feature extraction module may include a convolution layer, a pooling layer, and a fully connected layer. For example, for a text classification scenario, the convolution layer of the spatiotemporal feature extraction module may extract temporal features and / or spatial features through a convolution kernel, the pooling layer of the spatiotemporal feature extraction module may reduce feature dimensions through pooling, and the fully connected layer of the spatiotemporal feature extraction module may flatten the output features of the pooling layer and perform classification through the fully connected layer.

[0070] S203: The text feature extraction module extracts text features of the sample text data.

[0071] The text feature extraction module extracts the word vectors of each word in the sample text data and the relationship between the word vectors. For example, for the sample text data "I want to hire someone to repair my computer on Friday", the words "Friday", "repair" and "computer" in the sample text data "I want to hire someone to repair my computer on Friday" can be extracted, and the features of the words "Friday", "repair" and "computer" can be searched from the vocabulary feature library, and the features of the words "Friday", "repair" and "computer" are spliced, and the spliced ​​features obtained are used as the text features of the sample text data.

[0072] S204: The spatiotemporal feature extraction module extracts the spatiotemporal features of the sample text data.

[0073] The spatiotemporal feature extraction module extracts the word vectors of each word related to time and / or space in the sample text data and the relationship between the word vectors.

[0074] For example, for the sample text data "I want to hire someone to repair my computer on Friday", the word "Friday" related to time and / or space in the sample text data "I want to hire someone to repair my computer on Friday" can be extracted, and the features of the word "Friday" can be searched from the vocabulary feature library, and the features of the word "Friday" can be used as the spatiotemporal features of the sample text data.

[0075] For another example, for the sample text data "I want to hire someone to repair my computer on Friday", we can extract the time and / or space-related words "May" and "Shanghai" from the sample text data "I want to travel to Shanghai in May", and we can search for the features of the words "May" and "Shanghai" from the vocabulary feature library, and use the features of the words "May" and "Shanghai" as the spatiotemporal features of the sample text data.

[0076] In the present disclosure, there is no sequential execution order between S203 and S204.

[0077] S205, the model to be trained identifies the predicted intent corresponding to the sample text data based on the text features and the spatiotemporal features, and determines a first intent contribution rate corresponding to the text features and a second intent contribution rate corresponding to the spatiotemporal features.

[0078] In the present disclosure, the first intention contribution rate corresponding to the text feature is used to evaluate the importance of the text feature in determining the predicted intention. The second intention contribution rate corresponding to the spatiotemporal feature is used to evaluate the importance of the spatiotemporal feature in determining the predicted intention.

[0079] For example, the model to be trained can encode the text feature to obtain the text feature vector h_t, and linearly map the text feature vector h_t to obtain vector k and vector v. The model to be trained can also map the spatiotemporal feature h_v to obtain vector q. The model to be trained fuses vector k, vector v, and vector q to obtain a fused feature, identifies the intent corresponding to the fused feature, and obtains the predicted intent corresponding to the sample text data.

[0080] S206: Determine the loss function value of the model to be trained based on the label information and the predicted intent.

[0081] In the present disclosure, the loss function may adopt a logarithmic loss function, an exponential loss function, and the like.

[0082] S207: Determine whether the model to be trained has completed training based on the loss function value, the first intention contribution rate, and the second intention contribution rate.

[0083] In the present disclosure, a preset loss function threshold can be set, and the change trend of the loss function can be monitored by comparing the loss function value with the preset loss function threshold, and whether the model converges can be determined based on the change trend of the loss function.

[0084] The first intention contribution rate and the second intention contribution rate can reflect the importance of text features and spatiotemporal features. Whether to adjust the text feature extraction module and the spatiotemporal feature extraction module can be determined by the first intention contribution rate and the second intention contribution rate.

[0085] S208, determining the model obtained after completing the training as the target intent recognition model.

[0086] In one possible implementation, Figure 3 Another training flow diagram of the target intention recognition model provided in the embodiment of the present application is shown. Figure 3 As shown, the determining whether the to-be-trained model has completed training based on the loss function value, the first intention contribution rate, and the second intention contribution rate includes:

[0087] S301: If the loss function value is not less than a preset loss threshold, compare the first intention contribution rate and the second intention contribution rate.

[0088] In the present disclosure, the preset loss threshold can be set according to the scenario, for example, set to 0.05 or 0.1, etc. The loss function value is not less than the preset loss threshold, indicating that the model of the current model iteration has not met the training requirements, and the model parameters need to be adjusted and retrained if the number of iterations is not reached.

[0089] S302. If the contribution rate of the first intention is greater than that of the second intention, adjust the parameters of the text feature extraction module of the to-be-trained model.

[0090] If the contribution rate of the first intention is greater than that of the second intention, it indicates that the text features are more important for determining the predicted intention. Then, when the loss function value is not less than the preset loss threshold, the text features have a greater impact on the predicted intention determined by the model. Therefore, in order to improve the accuracy of the predicted intention, it is necessary to adjust the parameters of the text feature extraction module of the to-be-trained model.

[0091] S303. If the contribution rate of the second intention is greater than that of the first intention, adjust the parameters of the spatio-temporal feature extraction module of the to-be-trained model.

[0092] If the contribution rate of the second intention is greater than that of the first intention, it indicates that the spatio-temporal features are more important for determining the predicted intention. Then, when the loss function value is not less than the preset loss threshold, the spatio-temporal features have a greater impact on the predicted intention determined by the model. Therefore, in order to improve the accuracy of the predicted intention, it is necessary to adjust the parameters of the spatio-temporal feature extraction module of the to-be-trained model.

[0093] S304. For the new sample text data, return to execute the step of inputting the sample text data into the to-be-trained model until the loss function value is less than the preset loss threshold, and determine whether the to-be-trained model is completed.

[0094] In a possible implementation manner, determining the contribution rate of the first intention corresponding to the text features and the contribution rate of the second intention corresponding to the spatio-temporal features may include steps A1 - A4:

[0095] Step A1. Determine the first predicted classification result corresponding to the sample text data based on the text features.

[0096] In the present disclosure, the text feature extraction module may also perform intention prediction on the sample text data based on the text features, and the obtained predicted intention is used as the first predicted classification result.

[0097] Step A2. Determine the second predicted classification result corresponding to the sample text data according to the spatio-temporal features.

[0098] In the present disclosure, the spatio-temporal feature extraction module may also perform intention prediction on the sample text data based on the spatio-temporal features, and the obtained predicted intention is used as the second predicted classification result.

[0099] Step A3. Based on the predicted intention, determine the weight information of the first predicted classification result and the second predicted classification result.

[0100] In the present disclosure, the similarity between the first predicted classification result and the predicted intention can be calculated as the weight information of the first predicted classification result, and the similarity between the second predicted classification result and the predicted intention can be calculated as the weight information of the second predicted classification result. Alternatively, in the present disclosure, weights can also be assigned to the first predicted classification result and the second predicted classification result in advance according to the classification scenario.

[0101] Step A4: determining the first intention contribution rate corresponding to the text feature and the second intention contribution rate corresponding to the spatiotemporal feature respectively according to the weight of the first classification result and the weight of the second classification result.

[0102] In the present disclosure, the weight of the first classification result can be determined as the first intention contribution rate corresponding to the text feature, and the weight of the second classification result can be determined as the second intention contribution rate corresponding to the spatiotemporal feature. Alternatively, in the present disclosure, the first similarity between the first classification result and the true intention of the sample text data can be calculated based on the weight of the first classification result, and the second similarity between the second classification result and the true intention of the sample text data can be calculated based on the weight of the second classification result, and the first similarity can be determined as the first intention contribution rate corresponding to the text feature, and the second similarity can be determined as the second intention contribution rate corresponding to the spatiotemporal feature.

[0103] In one possible implementation, Figure 4 A schematic diagram of a structure of a model to be trained provided in an embodiment of the present application is shown. Figure 4 As shown, the model to be trained includes a text feature extraction module 401, a spatial feature extraction module 402 and a temporal feature extraction module 403. Figure 4 As shown in the figure, the sample text data "Country A imposes additional tariffs on new energy vehicles from Country B" is input into the model to be trained. The text feature extraction module of the model to be trained can encode the text data of "Country A imposes additional tariffs on new energy vehicles from Country B" and extract text features, and determine the label logits1 for the text feature. Figure 4 As shown, the spatial feature extraction module 402 can extract the spatial vocabulary "country B" from "country A imposes additional tariffs on new energy vehicles from country B", and perform spatial text encoding, perform feature extraction based on the encoded spatial text, and use the classification layer in combination with the extracted spatial features to predict the intention of "country A imposes additional tariffs on new energy vehicles from country B", and use the prediction result as the label logits2 of the spatial feature. The time feature extraction module 403 did not extract the time vocabulary from "country A imposes additional tariffs on new energy vehicles from country B". When extracting time features, the time feature extraction module 403 can perform time text encoding on the time features, perform feature extraction based on the encoded time text, and use the classification layer in combination with the extracted spatial features to predict the intention of the sample text data, and use the prediction result as the label logits3 of the time feature. As shown Figure 4 As shown, in the present disclosure, weights can also be assigned to various labels. For example, weight w1 is assigned to the label logits1 of the text feature, weight w2 is assigned to the label logits2 of the spatial feature, and weight w3 is assigned to the label logits3 of the time feature. The label label of the sample text data is determined by combining various labels and corresponding weights, and the predicted intent is determined based on the label label of the sample text data.

[0104] In a possible implementation, the target intent recognition model can also be updated based on the target intent and the target text. In the present disclosure, the target text can be used as training data, and the target intent can be used as label information of the training data to further update the target intent recognition model, so that the classification of the target intent recognition model can be more accurate and more generalized.

[0105] Based on the same inventive concept, according to the text classification method provided in the above embodiment of the present disclosure, correspondingly, another embodiment of the present disclosure further provides a text classification device, whose structural schematic diagram is shown in FIG. Figure 5 As shown, specifically including:

[0106] The data acquisition module 501 is used to acquire the text data to be recognized;

[0107] The data recognition module 502 is used to recognize the text data to be recognized by using the target intention recognition model, and extract the target text in the text data to be recognized, wherein the target text is a text associated with the temporal and spatial features;

[0108] Intention determination module 503, the user determines the target intention corresponding to the text data to be recognized under the specified time and space conditions based on the target text and the specified time and space conditions.

[0109] In one embodiment, the device further comprises:

[0110] A model training module is used to obtain sample text data and label information of the sample text data, wherein the label information represents the real intent corresponding to the sample text data; the sample text data is input into a model to be trained, wherein the model to be trained includes a text feature extraction module and a spatiotemporal feature extraction module; the text feature extraction module extracts the text features of the sample text data; the spatiotemporal feature extraction module extracts the spatiotemporal features of the sample text data; the model to be trained identifies the predicted intent corresponding to the sample text data based on the text features and the spatiotemporal features, and determines a first intent contribution rate corresponding to the text features and a second intent contribution rate corresponding to the spatiotemporal features; based on the label information and the predicted intent, determines a loss function value of the model to be trained; based on the loss function value, the first intent contribution rate and the second intent contribution rate, determines whether the model to be trained has completed training; and determines the model obtained after completing the training as a target intent recognition model.

[0111] In one possible implementation, the model training module is specifically used to compare the first intention contribution rate and the second intention contribution rate if the loss function value is not less than a preset loss threshold; if the first intention contribution rate is greater than the second intention contribution rate, adjust the parameters of the text feature extraction module of the model to be trained; if the second intention contribution rate is greater than the first intention contribution rate, adjust the parameters of the spatiotemporal feature extraction module of the model to be trained; for new sample text data, return to execute the step of inputting the sample text data into the model to be trained until the loss function value is less than the preset loss threshold, and determine whether the model to be trained has completed training.

[0112] In one possible implementation, the model training module is specifically used to determine a first predicted classification result corresponding to the sample text data based on the text feature; determine a second predicted classification result corresponding to the sample text data according to the spatiotemporal feature; determine weight information of the first predicted classification result and the second predicted classification result based on the predicted intent; and determine a first intent contribution rate corresponding to the text feature and a second intent contribution rate corresponding to the spatiotemporal feature according to the weight of the first classification result and the weight of the second classification result, respectively.

[0113] In one possible implementation, the specified spatiotemporal condition includes at least one of a time condition, a space condition, and a state condition.

[0114] In one possible implementation, the model training module is further used to update the target intent recognition model based on the target intent and the target text.

[0115] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.

[0116] Figure 6 A schematic block diagram of an example electronic device 400 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0117] like Figure 6 As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0118] A number of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0119] The computing unit 601 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as the text classification method. For example, in some embodiments, the text classification method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the text classification method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform the text classification method in any other appropriate manner (e.g., by means of firmware).

[0120] The electronic device 600 may further include an image acquisition device.

[0121] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), integrated systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0122] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0123] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0124] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0125] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0126] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0127] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0128] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of the present disclosure, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0129] The above is only a specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art who is familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present disclosure, which should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be based on the protection scope of the claims.

Claims

1. A text classification method, the method comprising: Obtaining text data to be recognized; Using a target intention recognition model to recognize the text data to be recognized, and extracting a target text from the text data to be recognized, wherein the target text is a text associated with a temporal and spatial feature; Based on the target text and the specified time and space conditions, the target intent corresponding to the text data to be recognized under the specified time and space conditions is determined.

2. According to the method of claim 1, the training scheme of the target intent recognition model comprises: Acquire sample text data and label information of the sample text data, wherein the label information represents the real intention corresponding to the sample text data; Inputting the sample text data into a model to be trained, wherein the model to be trained includes a text feature extraction module and a spatiotemporal feature extraction module; The text feature extraction module extracts text features of the sample text data; The spatiotemporal feature extraction module extracts the spatiotemporal features of the sample text data; The model to be trained identifies the predicted intent corresponding to the sample text data based on the text features and the spatiotemporal features, and determines a first intent contribution rate corresponding to the text features and a second intent contribution rate corresponding to the spatiotemporal features; Based on the label information and the predicted intent, determining a loss function value of the model to be trained; Determining whether the model to be trained has completed training based on the loss function value, the first intention contribution rate, and the second intention contribution rate; The model obtained after training is determined as the target intent recognition model.

3. The method according to claim 2, wherein determining whether the model to be trained has completed training based on the loss function value, the first intention contribution rate, and the second intention contribution rate comprises: If the loss function value is not less than a preset loss threshold, comparing the first intention contribution rate and the second intention contribution rate; If the first intention contribution rate is greater than the second intention contribution rate, adjusting the parameters of the text feature extraction module of the model to be trained; If the second intention contribution rate is greater than the first intention contribution rate, adjusting the parameters of the spatiotemporal feature extraction module of the model to be trained; For new sample text data, return to the step of inputting the sample text data into the model to be trained until the loss function value is less than the preset loss threshold, and determine whether the model to be trained has completed training.

4. According to the method of claim 2, the determining the first intention contribution rate corresponding to the text feature and the second intention contribution rate corresponding to the spatiotemporal feature comprises: Determine a first prediction classification result corresponding to the sample text data based on the text feature; Determine a second prediction classification result corresponding to the sample text data according to the spatiotemporal feature; Based on the prediction intention, determining weight information of the first prediction classification result and the second prediction classification result; According to the weight of the first classification result and the weight of the second classification result, a first intention contribution rate corresponding to the text feature and a second intention contribution rate corresponding to the spatiotemporal feature are determined respectively.

5. According to the method of claim 3, the specified spatiotemporal condition comprises at least one of a time condition, a space condition and a state condition.

6. The method according to claim 1, further comprising: The target intent recognition model is updated based on the target intent and the target text.

7. A text classification device, comprising: A data acquisition module, used to acquire text data to be recognized; A data recognition module, used to recognize the text data to be recognized by using a target intention recognition model, and extract a target text from the text data to be recognized, wherein the target text is a text associated with a temporal and spatial feature; The intention determination module is used for determining the target intention corresponding to the text data to be identified under the specified time and space conditions based on the target text and the specified time and space conditions.

8. The device according to claim 7, further comprising: A model training module, used to obtain sample text data and label information of the sample text data, wherein the label information represents the real intention corresponding to the sample text data; The sample text data is input into a model to be trained, and the model to be trained includes a text feature extraction module and a spatiotemporal feature extraction module; the text feature extraction module extracts text features of the sample text data; the spatiotemporal feature extraction module extracts spatiotemporal features of the sample text data; the model to be trained identifies the predicted intent corresponding to the sample text data based on the text features and the spatiotemporal features, and determines a first intent contribution rate corresponding to the text features and a second intent contribution rate corresponding to the spatiotemporal features; based on the label information and the predicted intent, determines a loss function value of the model to be trained; Determining whether the model to be trained has completed training based on the loss function value, the first intention contribution rate, and the second intention contribution rate; The model obtained after training is determined as the target intent recognition model.

9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 6 is implemented.

10. A storage medium comprising computer executable instructions, which when executed by a computer processor are used to perform the method of any one of claims 1 to 6.