Text processing method and related device

By combining semantic recognition and attribute information of call text, the conversation text that does not belong to the main conversation object can be accurately identified and stripped, which solves the efficiency and accuracy problems of recognition and stripping in the existing technology and realizes efficient text processing.

CN120688505APending Publication Date: 2025-09-23MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510080764.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

It is difficult for existing technologies to efficiently and accurately identify the main and secondary conversation participants in call texts and to remove text that does not belong to the main conversation participant.

Method used

By performing semantic recognition on the first conversation text of the first conversation object and combining the attribute information of the first conversation object, a second conversation sentence belonging to the object in the first conversation sentence is determined, and the first conversation text is updated based on the second conversation sentence.

Benefits of technology

The recognition accuracy of conversational sentences and the efficiency of text processing are improved. Only the recognized conversational sentences are judged, which reduces the probability of misjudgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688505A_ABST
    Figure CN120688505A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a text processing method and a related device, which are at least applied to the technical field of computers, and the text processing method comprises the following steps: performing semantic recognition on a first session text of a first session object, and determining a first session sentence in the first session text, the first session sentence comprising session sentences of a plurality of session objects; according to the attribute information of the first session object, determining a second session sentence belonging to the first session object in the first session sentence; and updating the first session text based on the second session sentence to obtain a second session text belonging to the first session object. After the first session sentence in the first session text is recognized, the second session text belonging to the first session object in the first session sentence is judged in combination with the attribute information of the first session object, and the session text not belonging to the first session object can be accurately stripped.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and are related to, but not limited to, a text processing method and related devices. Background Art

[0002] In some specific scenarios, it's necessary to perform semantic recognition on the conversation text generated during a call to identify the primary and secondary participants in the text and remove text that doesn't belong to the primary participant. Related technologies often rely on simple sentence segmentation of the conversation text. However, when the text contains semantically confusing sentences, it's not possible to efficiently and accurately identify the primary and secondary participants in the conversation text and precisely remove the text corresponding to the primary and secondary participants. Summary of the Invention

[0003] The embodiments of the present application provide a text processing method and related apparatus, which can be applied at least in the field of computer technology. By identifying a first conversation sentence in a first conversation text and combining it with attribute information of a first conversation object, a second conversation text belonging to the first conversation object in the first conversation sentence can be judged, and conversation text that does not belong to the first conversation object can be accurately removed.

[0004] The technical solution of the embodiment of the present application is implemented as follows:

[0005] An embodiment of the present application provides a text processing method, comprising: performing semantic recognition on a first conversation text of a first conversation object, determining a first conversation sentence in the first conversation text, wherein the first conversation sentence includes conversation sentences of multiple conversation objects; determining a second conversation sentence in the first conversation sentence belonging to the first conversation object based on attribute information of the first conversation object; and updating the first conversation text based on the second conversation sentence to obtain a second conversation text belonging to the first conversation object.

[0006] An embodiment of the present application provides a text processing device, comprising: a first determination module configured to perform semantic recognition on a first conversation text of a first conversation object, and determine a first conversation sentence in the first conversation text, wherein the first conversation sentence includes conversation sentences of multiple conversation objects; a second determination module configured to determine, based on attribute information of the first conversation object, a second conversation sentence in the first conversation sentence that belongs to the first conversation object; and an update module configured to update the first conversation text based on the second conversation sentence to obtain a second conversation text belonging to the first conversation object.

[0007] An embodiment of the present application provides an electronic device, comprising: a memory for storing executable instructions; and a processor for implementing the above-mentioned text processing method when executing the executable instructions stored in the memory.

[0008] An embodiment of the present application provides a computer program product, which includes a computer program or executable instructions. When the computer program or computer executable instructions are executed by a processor, the text processing method provided by the embodiment of the present application is implemented.

[0009] An embodiment of the present application provides a computer-readable storage medium storing a computer program or executable instructions for implementing the text processing method provided in the embodiment of the present application when executed by a processor.

[0010] The above solution has the following beneficial effects:

[0011] After identifying the first conversation sentence in the first conversation text, the embodiment of the present application combines the attribute information of the first conversation object to determine the second conversation sentence belonging to the first conversation object in the first conversation sentence, and updates the first conversation text based on the second conversation sentence. In this way, since the second conversation sentence belonging to the first conversation object is determined in the first conversation sentence based on the attribute information of the first conversation object, the recognition accuracy of the second conversation sentence can be improved. Moreover, the judgment is only performed on the identified first conversation sentence, rather than on all sentences in the first conversation text, thereby improving the efficiency of text processing. It can be seen that the text processing method provided by the embodiment of the present application can not only accurately identify the second conversation sentence belonging to the first conversation object in the first conversation sentence, but also improve the efficiency of text processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 This is an optional architectural diagram of a text processing system provided in an embodiment of the present application;

[0013] Figure 2 This is an optional flowchart of the text processing method provided in the embodiment of the present application;

[0014] Figure 3 This is another optional flowchart of the text processing method provided in the embodiment of the present application;

[0015] Figure 4 This is a flow chart of performing text language category conversion processing on a third conversation text to obtain a first conversation text of a first type of language type, as provided in an embodiment of the present application;

[0016] Figure 5 This is a schematic diagram of a process for determining a first conversation statement provided in an embodiment of the present application;

[0017] Figure 6 This is a flowchart of determining a first conversation sentence from a first conversation text based on a third sentence provided by an embodiment of the present application;

[0018] Figure 7 This is a flowchart of determining a second conversation statement based on attribute information of a first conversation object provided by an embodiment of the present application;

[0019] Figure 8 This is a text processing framework diagram provided by an embodiment of the present application;

[0020] Figure 9 is a structural diagram of a text processing device provided in an embodiment of the present application;

[0021] Figure 10 It is a schematic diagram of the composition structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0023] In the following description, reference is made to "some embodiments," which describe a subset of all possible embodiments. However, it will be understood that "some embodiments" may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict. Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by those skilled in the art to which the embodiments of this application pertain. The terms used in the embodiments of this application are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0024] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0025] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0026] Before describing the text processing method provided in the embodiments of the present application, the professional terms involved in the embodiments of the present application are first described:

[0027] 1) Semantically mixed sentences: refers to sentences in a conversation text that contain multiple objects. That is, semantically mixed sentences are sentences that easily lead to confusing semantic expressions and unclear semantics. The multiple objects include the conversation object corresponding to the conversation text and other conversation objects.

[0028] 2) Clustering: A method of grouping data objects based on their similarity or proximity to form different classes or clusters. This allows data objects within the same cluster to be highly similar under a certain metric, while samples in different clusters are more different.

[0029] 3) Machine Learning: It is a technology that uses data-driven algorithms and models to enable computers to automatically learn and improve task performance without explicit programming instructions.

[0030] Before explaining the text processing method of the embodiment of the present application, here, we first explain the exemplary application of the text processing device of the embodiment of the present application, which is an electronic device for implementing the text processing method. In one implementation, the text processing device (i.e., electronic device) provided in the embodiment of the present application can be implemented as a terminal or as a server. In one implementation, the electronic device provided in the embodiment of the present application can be implemented as any terminal with text processing function, such as a laptop computer, a tablet computer, a desktop computer, an intelligent robot, a smart home appliance, and an intelligent vehicle-mounted device; in another implementation, the text processing device provided in the embodiment of the present application can also be implemented as a server, wherein the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDNs), and big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiment of the present application. Below, we will explain the exemplary application of the text processing device when it is implemented as a server.

[0031] See also Figure 1 , Figure 1This is an optional architectural diagram of a text processing system provided in an embodiment of the present application. The text processing system 10 in this embodiment of the present application includes at least a terminal 100, a network 200, and a server 300. A text processing application is deployed on the terminal 100, and the server 300 can be the background server of the text processing application. The server 300 can constitute the text processing device in this embodiment of the present application, that is, the text processing method in this embodiment of the present application is implemented through the server 300. The terminal 100 is connected to the server 300 via the network 200. The network 200 can be a wide area network or a local area network, or a combination of the two.

[0032] See also Figure 1 Terminal 100 receives a text processing operation from a user and generates a text processing request in response to the text processing operation. Terminal 100 then sends the text processing request to server 300 via network 200. In response to the text processing request, server 300 obtains a first conversation text of a first conversation object and performs semantic recognition on the first conversation text of the first conversation object to determine a first conversation sentence in the first conversation text, where the first conversation sentence includes conversation sentences of multiple conversation objects. Next, server 300 determines a second conversation sentence in the first conversation sentence that belongs to the first conversation object based on attribute information of the first conversation object. Next, server 300 updates the first conversation text based on the second conversation sentence to obtain a second conversation text belonging to the first conversation object, and determines the second conversation text as the text processing result. Finally, server 300 sends the text processing result to terminal 100 via network 200 and displays the text processing result on a display interface of terminal 100.

[0033] In some embodiments, the above-mentioned text processing method can also be executed by the terminal, that is, after receiving the user's text processing operation, the terminal 100 can respond to the text processing operation by obtaining the first conversation text of the first conversation object, performing semantic recognition on the first conversation text of the first conversation object, and determining the first conversation sentence in the first conversation text, wherein the first conversation sentence includes conversation sentences of multiple conversation objects; then, the terminal 100 determines the second conversation sentence belonging to the first conversation object in the first conversation sentence based on the attribute information of the first conversation object; then, the terminal 100 updates the first conversation text based on the second conversation sentence to obtain the second conversation text belonging to the first conversation object, and determines the second conversation text as the text processing result; finally, the terminal 100 displays the text processing result on the display interface.

[0034] The text processing methods provided in each embodiment of the present application can be executed by an electronic device, wherein the electronic device can be a server or a terminal, that is, the text processing methods in each embodiment of the present application can be executed by a server, or by a terminal, or by interaction between a server and a terminal.

[0035] Figure 2 This is an optional flow chart of the text processing method provided in the embodiment of the present application. Figure 2 The steps shown are explained as Figure 2 As shown, the text processing method is described by taking the execution subject of the server as an example. The method includes the following steps S101 to S103:

[0036] Step S101 : performing semantic recognition on a first conversation text of a first conversation object, and determining a first conversation sentence in the first conversation text, wherein the first conversation sentence includes conversation sentences of a plurality of conversation objects.

[0037] The first conversation object refers to the main object participating in the interaction in the conversation. For example, user A is on the phone with user B, and user C is talking next to user B. In this case, user B is the first conversation object.

[0038] The first conversation text refers to a specific text segment that needs to be processed or analyzed. The first conversation text can be text data obtained by the user online, text data obtained from an existing text database, or text data generated by the user by converting a voice file or video file. This application has no restrictions on the size of the first conversation text; it can be a sentence, a paragraph, or the entire conversation content. The first conversation text can include simple declarative sentences, complex compound sentences, and even paragraphs composed of multiple sentences to meet different needs.

[0039] Semantic recognition generally refers to extracting useful information from the text to be processed, including specific sentence types, keywords, entity words, etc. In the embodiment of the present application, semantic recognition refers to the process of determining the first conversation sentence in the first conversation text. Semantic recognition can be achieved through machine learning algorithms, such as k-means clustering algorithm, logistic regression, etc.; it can also be achieved through deep learning algorithms, such as recurrent neural networks, large language models, etc.; it can also be achieved by combining machine learning algorithms and deep learning algorithms, such as combining the bidirectional transformer model (BERT, Bidirectional Encoder Representations from Transformers) and the k-means clustering algorithm.

[0040] A first conversation sentence refers to a sentence in a conversation text that contains conversation sentences of multiple conversation objects, that is, the first conversation sentence is a sentence that easily leads to confusion in semantic expression and unclear semantics; multiple conversation objects include conversation objects corresponding to the first conversation text and other conversation objects; for example, staff member A is having a voice call with user B, and user B's colleague user C is next to user B. During the call between staff member A and user B, user C has also been talking to other people. After the call ends, the call recording is converted into text data, and the text data contains the conversation texts corresponding to staff member A, user B and user C. At this time, user B is the conversation object corresponding to the text data, and user C is another object. Since what user C said has nothing to do with the conversation between staff member A and user B, the semantics of some sentences in the generated text data will be unclear. These semantically unclear sentences are the first conversation sentences.

[0041] In some embodiments, step S101 can be implemented by the following method: first, the server divides the first conversation text into sentences using a word segmentation tool to obtain a first sentence set; then, the sentences in the first sentence set are clustered using a k-means clustering algorithm to obtain a first clustering result, wherein the first clustering result includes multiple first cluster sets, and each first cluster set contains a first central sentence and at least one second sentence, the first central sentence refers to the sentence that best represents the core expression of the corresponding first cluster set, and the second sentence is the sentence belonging to each first cluster set other than the first central sentence, such as: the first cluster set 1 contains sentences 1 to 5, wherein sentence 3 is the first central sentence, and sentences 1, 2, 4, and 5 are second sentences. The specific clustering algorithm is not limited in this application.

[0042] Next, the second sentence whose distance from the first central sentence is greater than a first preset threshold is determined as the third sentence; finally, the first conversation sentence is determined based on the third sentence; wherein, when determining the first conversation sentence based on the third sentence, there are two situations:

[0043] Case 1: Assume that a clustering operation is performed on the statements in the first statement set to obtain 10 first cluster sets, among which the third statement A exists in the first first cluster set and the third statement B exists in the second first cluster set. Then, the second statements other than the third statement A and the third statement B in the first and second first cluster sets are deleted, and the third statement A and the third statement B are combined with all the statements in the third to tenth first cluster sets to form a third statement set. Then, a clustering operation is performed on the third statement set to obtain a second clustering result. For example, the second clustering result includes 8 second cluster sets, among which the first second cluster set contains the third statement A and the second second cluster set contains the third statement B. When the distance between the third statement A and the central sentence of the first second cluster set is greater than a second preset threshold, the third statement A is determined to be the first conversation sentence. When the distance between the third statement B and the central sentence of the second second cluster set is less than the second preset threshold, it is determined that the third statement B is not the first conversation sentence.

[0044] Case 2: Assume that a clustering operation is performed on the statements in the first statement set, resulting in 10 first clusters. The first first cluster contains third statement A, and the second first cluster contains third statement B. First, all second statements in the first first cluster except third statement A are deleted, and third statement A and all statements in the second through tenth first clusters are combined to form a third statement set. Then, a clustering operation is performed on the third statement set to obtain a second clustering result. For example, if the second clustering result includes 8 second clusters, of which the first second cluster contains third statement A, when the distance between third statement A and the central sentence of the first second cluster is greater than a second preset threshold, third statement A is determined to be a first conversation statement. While determining third statement A, the same operation steps are performed on third statement B to determine whether third statement B is a first conversation statement.

[0045] It should be noted that identifying the first conversation sentence in the first conversation text in step S101 can provide effective data support for subsequent text processing, thereby improving the efficiency and quality of text processing.

[0046] Step S102: determining, according to the attribute information of the first conversation object, a second conversation statement in the first conversation statement that belongs to the first conversation object.

[0047] The attribute information of the first conversation partner refers to information describing the various attributes and characteristics of the first conversation partner. This attribute information can be either explicit or implicit. Explicit attributes include the first conversation partner's name, age, gender, and occupation, while implicit attributes include the first conversation partner's current emotional state (e.g., anger, joy) and conversational style (e.g., formal, casual). Users can directly enter the attribute information of the first conversation partner, or extract it from the first conversation text using natural language processing, machine learning, or deep learning.

[0048] The second conversation sentence refers to a sentence belonging to the first conversation object that is identified in the first conversation sentence based on the attribute information of the first conversation object.

[0049] In some embodiments, step S102 may be implemented by the following method: first, the server uses a natural language processing tool to identify major punctuation points in the first conversation sentence, such as punctuation marks and conjunctions. Then, based on the identification results, the first conversation sentence is segmented into multiple sentence segments. For a complex first conversation sentence, semantic analysis may be combined to identify the core semantics of different segments in the first conversation sentence, and the first conversation sentence is segmented according to the core semantics. The core semantics refers to important concepts, themes, or key information in a sentence. For example, keyword extraction is used to identify multiple keywords in the first conversation sentence, and semantic boundaries are determined based on the association between the words. After the keywords and semantic boundaries are determined, the first conversation sentence is segmented. For example, if the first conversation sentence is "I work at home today, let's eat rice tomorrow. The weather suddenly turned cold, and I will go home soon", keyword extraction obtains the keywords "work at home", "rice", and "go home", semantic boundaries are determined based on the association between each word in the first conversation sentence, and the first conversation sentence is segmented into three sentence segments, namely "I work at home today", "let's eat rice tomorrow", and "The weather suddenly turned cold, and I will go home soon".

[0050] Next, feature extraction is performed on the attribute information of the first conversation object and multiple sentence fragments respectively, and an object feature vector and multiple sentence fragment vectors are obtained accordingly; then, the object feature vector is concatenated with each sentence fragment vector in the multiple sentence fragment vectors to obtain a concatenated vector, for example, if the sentence fragment vector is [a, b, c] and the object feature vector is [e, f], then the concatenated vector is [a, b, c, e, f]; then, the concatenated vector is input into the classification model, and a probability value is output, which represents the probability that the sentence fragment belongs to the first conversation object; finally, the relationship between the probability value output for each sentence fragment and a preset probability threshold is determined. When the output probability value is greater than the preset probability threshold, it is determined that the sentence fragment belongs to the first conversation object.

[0051] It should be noted that step S102 can more accurately identify conversation sentences that do not belong to the first conversation object in the first conversation sentences by combining the attribute information of the first conversation object, thereby improving recognition accuracy and reducing the probability of misjudgment.

[0052] Step S103: updating the first conversation text based on the second conversation statement to obtain a second conversation text belonging to the first conversation object.

[0053] In some embodiments, step S103 can be implemented by the following method: after determining the second conversation statement belonging to the first conversation object in the first conversation statement, use the second conversation statement to replace the first conversation statement in the first conversation text, thereby updating the first conversation text and obtaining the second conversation text belonging to the first conversation object.

[0054] The text processing method of the embodiment of the present application can provide effective data support for subsequent text processing by identifying the first conversation sentence in the first conversation text, thereby improving the efficiency and quality of text processing. Moreover, by combining the attribute information of the first conversation object, the text belonging to the first conversation object can be more accurately identified in the first conversation sentence, thereby improving the recognition accuracy and reducing the probability of misjudgment, thereby efficiently and accurately removing the text that does not belong to the first conversation object in the first conversation sentence.

[0055] The following examples illustrate application scenarios of the text processing method provided in the embodiments of the present application. The embodiments of the present application can be applied to at least the following exemplary scenarios:

[0056] In a scenario of bank resource recovery, staff member A conducts a voice call with user B. However, during the call, there are other users speaking around user B. After the communication is completed, staff member A needs to remove the voice texts of other users from the call text generated during the communication process, retaining only the voice text of user B. In this way, the text processing method provided by the embodiment of the present application can be adopted. First, the terminal generates a text processing request in response to staff member A's text processing operation and sends the text processing request to the server. Then, after receiving the text processing request, the server responds to the text processing request, obtains the first conversation text of the first conversation object (user B), and performs semantic recognition on the first conversation text of the first conversation object to determine the first conversation sentence in the first conversation text, wherein the first conversation sentence includes conversation sentences of multiple conversation objects (user B and its characteristic speaking user); then, based on the attribute information of the first conversation object, the second conversation sentence belonging to the first conversation object in the first conversation sentence is determined; finally, the first conversation text is updated based on the second conversation sentence to obtain the second conversation text belonging to the first conversation object.

[0057] Based on the above scenario, the text processing method of the embodiment of the present application is explained. Figure 3 This is another optional flow chart of the text processing method provided in the embodiment of the present application, such as Figure 3 As shown, the method includes the following steps S201 to S210:

[0058] Step S201: The terminal receives a text processing operation.

[0059] Here, a text processing application may be running on the terminal, and the server constitutes a background server of the text processing application. The text processing operation may be a selection operation or an input operation input by a user through a client of the text processing application running on the terminal. For example, the user may select a first conversation text or a conversation voice for text processing, or may input the first conversation text or a conversation voice for text processing on the client.

[0060] In some embodiments, the client of the text processing application may provide an input interface or input box, allowing the user to select or input the first conversation text or conversation voice for text processing. The input interface may provide selection or input options in the form of a form, a text box, or a drop-down menu, and the specific form is not limited in this application. The user can select the first conversation text or conversation voice for text processing from predetermined options, or manually input the first conversation text or conversation voice for text processing.

[0061] Step S202: The terminal generates a text processing request in response to the text processing operation.

[0062] Here, the terminal may encapsulate the first conversation text or conversation voice selected or input by the user for text processing into a text processing request.

[0063] In some embodiments, in order to ensure the security of text processing requests, authentication parameters are added to the text processing requests. The authentication parameters are used to verify the legitimacy of the text processing requests to ensure that only authorized users can access the text processing services and to prevent access and abuse by unauthorized users. A common authentication method is to use an API key or token. An API key is a unique string of characters used to identify and verify the identity of a user. A token is a credential similar to an access token or authentication token that contains the user's identity information and permissions.

[0064] Step S203: The terminal sends a text processing request to the server.

[0065] In some embodiments, the terminal sends the encapsulated text processing request to the server and requests the server to perform text processing operations, usually using protocols such as HyperText Transfer Protocol (HTTP) or Web Socket to send the text processing request.

[0066] Step S204: The server obtains the first conversation text of the first conversation object in response to the text processing request.

[0067] Here, after receiving a text processing request, the server parses the text processing request. For example, for an HTTP request, the server may parse the request header and request body. The server may parse the request header to obtain relevant information about the request and parse the request body to obtain the main body data of the request, namely, the first conversation text or conversation voice for text processing. The server may parse specific fields or parameters in the request body, which contain the first conversation text or conversation voice for text processing, extract a specific data format from the request body, such as a lightweight data exchange format (JSON, JavaScript Object Notation) or Extensible Markup Language (XML), and then parse the data format to obtain the first conversation text or conversation voice for text processing.

[0068] In some embodiments, when the input is conversational speech, the conversational speech needs to be converted into text to obtain a first conversational text. Specifically, converting the conversational speech into text to obtain the first conversational text can be achieved by the following method: first, recognizing the conversational speech of the first conversation partner to obtain the language type of the conversational speech; then, if the language type is the first type, converting the conversational speech into the first conversational text using the first model; finally, if the language type is the second type, converting the conversational speech into a third conversational text of the second type using the first model; and performing text language category conversion processing on the third conversational text to obtain the first conversational text of the first type.

[0069] Here, the first type may be standard Mandarin; the second type may be other natural languages ​​that are different from the first type, such as dialects or foreign languages; the language type refers to the type of natural language, such as Chinese, English, French, etc. Chinese can also include Mandarin and dialects.

[0070] The first model refers to a model used to convert a speech signal into corresponding text data. The first model can be implemented as a deep neural network model, a long short-term memory network, etc. The specific model is not limited in this application.

[0071] The language category conversion process refers to the process of converting one language category into another, such as the process of converting a dialect into Mandarin. The language category conversion process can be performed through a preset conversion rule library or through an automatic conversion using deep learning algorithms. The deep learning algorithms can be implemented as convolutional neural networks, recurrent neural networks, etc. The specific conversion method and the specific deep learning algorithm are not limited in this application.

[0072] As an example of step S204, when the input data is session speech, the server extracts the time-frequency features of the session speech through a convolutional neural network. The time-frequency features can effectively represent the acoustic feature differences of different languages, thereby identifying the language type of the session speech. Then, when the language type is Mandarin, the server converts the session speech into a first session text through the first model. When the language type is a dialect or English, the server converts the session speech into a third session text through the first model, and performs lexical conversion or text translation on the third session text to convert the second session text into the first session text.

[0073] It should be noted that the embodiments of this application can identify and process speech inputs of different language types, thereby adapting to different application scenarios, meeting diverse user needs, and can be extended to more languages and text formats, enhancing the scalability and application scope of the system; moreover, unifying all texts into the same language or format facilitates subsequent processing and analysis, thereby improving the accuracy and efficiency of text processing.

[0074] In some embodiments, refer to Figure 4 , Figure 4 is a schematic flowchart of the process of performing text language category conversion processing on the third session text to obtain a first session text with a language type of the first type provided by the embodiments of this application. Figure 4 It shows that the server can perform text language category conversion processing on the third session text to obtain a first session text with a language type of the first type through the following steps S301 to step S303:

[0075] Step S301, the server replaces the second type of vocabulary in the third session text with the corresponding first type of vocabulary according to the mapping relationship between the preset second type of vocabulary and the first type of vocabulary, to obtain a session conversion text.

[0076] The second type of vocabulary refers to the special vocabulary in the second type of language that is different from Mandarin, such as the vocabulary in English or a dialect that is different from Mandarin. For example, if the second type of vocabulary is "rare", which is not said like this in Mandarin, then "rare" is the special vocabulary that is different from Mandarin.

[0077] The first type of vocabulary refers to the Mandarin vocabulary corresponding to the second type of vocabulary; for example, if the second type of vocabulary is "Apple", the first type of vocabulary corresponding to this second type of vocabulary is "苹果". The mapping relationship refers to the predefined corresponding relationship between two types or multiple types of vocabulary.

[0078] In some embodiments, step S301 can be implemented by the following method: First, when the server obtains the third session text under the second type, it traverses each vocabulary in the third session text and looks up the corresponding first type of vocabulary in the preset mapping table. If the first type of vocabulary corresponding to the second type of vocabulary is found, the second type of vocabulary is replaced with the corresponding first type of vocabulary. If a certain second type of vocabulary does not have a corresponding first type of vocabulary in the preset mapping table, the original second type of vocabulary can be selected and marked (such as bold or highlighted) for the second type of vocabulary without a corresponding first type of vocabulary; then, after all the vocabulary conversions are completed, the session conversion text is obtained.

[0079] As an example of step S301, if the third session text is "俺稀罕你", the server traverses each vocabulary in the third session text and looks up the first type of vocabulary corresponding to the second type of vocabulary in the preset mapping table. Among them, the second type of vocabulary "俺" corresponds to the first type of vocabulary "我", and the second type of vocabulary "稀罕" corresponds to the first type of vocabulary "喜欢". After finding the corresponding first type of vocabulary, the first type of vocabulary is used to replace the second type of vocabulary to obtain the session conversion text. Then, the session conversion text corresponding to the third session text "俺稀罕你" is "我喜欢你".

[0080] Step S302, the server standardizes the sentence structure in the session conversion text according to the predefined sentence conversion rules to obtain the standardized session text.

[0081] The predefined sentence conversion rules refer to the grammar and sentence adjustment rules set in advance, which are used to standardize and unified the sentence structure of the text.

[0082] The standardization process refers to unifying different forms of sentences into a predefined standard form. Through the standardization process, the consistency of the grammar and logical structure of the text can be ensured.

[0083] In some embodiments, step S302 can be implemented by the following method: First, the server analyzes each sentence in the session conversion text and identifies the grammar structure of each sentence (such as subject, predicate, object, etc.); then, according to the predefined sentence conversion rules, the sentences that meet the conversion rules are reorganized, replaced or simplified to ensure that the grammar structure of the sentences meets the standardization requirements, and the standardized session text is obtained.

[0084] As an example of step S302, if the conversation conversion text is "Have you eaten?", the server performs grammatical structure recognition on the conversation conversion text and determines that the conversation conversion text is an inverted sentence; then, according to the predefined sentence conversion rules, the conversion text is reorganized according to the sentence structure of subject + predicate + object to obtain a standardized conversation text. The standardized conversation text corresponding to the conversation conversion text is "Have you eaten?"

[0085] Step S303: The server determines the standardized conversation text as the first conversation text.

[0086] After obtaining the standardized processed text, the server determines the standardized conversation text as the first conversation text.

[0087] In step S205 , the server performs semantic recognition on the first conversation text of the first conversation object, and determines a first conversation sentence in the first conversation text, where the first conversation sentence includes conversation sentences of multiple conversation objects.

[0088] In some embodiments, see Figure 5 , Figure 5 This is a flowchart of determining a first conversation statement provided in an embodiment of the present application. Figure 5 It is shown that in step S205, the server performs semantic recognition on the first conversation text of the first conversation object and determines the first conversation sentence in the first conversation text, which can be achieved by the following steps S2051 to S2054:

[0089] Step S2051: The server divides the first conversation text into sentences to obtain a first sentence set.

[0090] In some embodiments, the server may use regular expressions, word segmentation tools, or machine learning algorithms to identify and extract sentence boundaries, such as identifying punctuation marks or other features (such as line breaks, indentations, etc.) in the first conversation text, and using the identified punctuation marks or other features as signs of sentence division; after the statement division is completed, the divided sentences are stored in a set data structure to obtain a first sentence set.

[0091] In step S2052, the server performs a clustering operation on the statements in the first statement set to obtain a first clustering result; each first clustering set of the first clustering result includes a first central sentence and K second statements; K is a positive integer.

[0092] The second statement refers to the other statements except the first central sentence that are divided into the first cluster set; for example, there are 50 statements in the first statement set, and 5 first cluster sets are obtained after the clustering operation, among which statements 1 to 9 belong to the first first cluster set, and statement 3 is the central sentence in the first first cluster set, that is, the first central sentence, then statements 1, 2, 4-9 are the second statements.

[0093] The clustering operation refers to the process of grouping the sentences in the first sentence set according to a certain similarity standard (such as semantic similarity, topic similarity, structural similarity, etc.).

[0094] In some embodiments, the clustering operation can be implemented by the following method: first, the server needs to extract features from each statement in the first statement set to obtain a feature vector corresponding to each statement; then, the similarity measurement between the statements is calculated based on the feature vector corresponding to each statement. Commonly used similarity measurement methods include Euclidean distance, cosine similarity, etc.; finally, the server uses a clustering algorithm based on the similarity measurement between each statement to group the statements in the first statement set and divide statements with higher similarity into one category.

[0095] After classifying the sentences, it is necessary to determine the first central sentence in each cluster set. The first central sentence can be determined by a centroid-based method, a density-based method, or a hierarchical clustering-based method. The specific determination method is not limited in this application. Take the centroid-based method as an example: first, average the feature vectors of all sentences belonging to the same cluster set to obtain the centroid of the class. The centroid does not necessarily correspond to an actual sentence, but is an abstract average vector; then, calculate the distance between each sentence in the class and the centroid (the commonly used distance is Euclidean distance, cosine similarity), and select the sentence closest to the centroid as the central sentence.

[0096] In step S2053 , the server determines a third sentence in each cluster set from the K second sentences based on the distance between the second sentence and the first central sentence, where the distance between the third sentence and the first central sentence is greater than a first preset distance threshold.

[0097] The distance between the second sentence and the first central sentence refers to mapping the second sentence and the first central sentence into the vector space to obtain the corresponding sentence vectors. The Euclidean distance between the two sentence vectors is the distance between the two sentences. The third sentence refers to the sentence among the K second sentences whose distance from the first central sentence is greater than the preset threshold.

[0098] As an example of step S2053, when there are 10 first cluster sets, the server will determine the distance between the second statement in each first cluster set and the first central sentence in the corresponding first cluster set. When the distance between the second statement A in the first first cluster set and the first central sentence in the first first cluster set is greater than the first preset distance threshold, the second statement A will be determined as the third statement in the first first cluster set.

[0099] In some embodiments, the preset distance threshold can be determined by the following method: first, manually annotate the intent of each text in a preset text dataset, where the preset text dataset includes semantic single sentences and mixed conversation sentences; then, obtain the sentence vector of each text in the preset text dataset through a pre-trained vector extraction model, and calculate the sentence vector distance between any two sentence vectors; then, calculate the mean of the farthest sentence vector distances of the semantic single sentences belonging to the same intent, and the mean of the nearest sentence vector distances of the mixed conversation sentences belonging to the same intent; finally, average the mean of the farthest sentence vector distances and the mean of the nearest sentence vector distances to obtain the preset distance threshold.

[0100] For example: the preset text data set is divided into 5 types of intents, namely Figure 1 ,meaning Figure 2 ,meaning Figure 3 ,meaning Figure 4 Harmony Figure 5 ;care Figure 1 The furthest vector distance of a single semantic sentence is 0.9, and the closest vector distance of a mixed conversation sentence is 0.3. Figure 2 The furthest vector distance of a single semantic sentence is 0.88, and the closest vector distance of a mixed conversation sentence is 0.2. Figure 3 The furthest vector distance of a single semantic sentence is 0.85, and the closest vector distance of a mixed conversation sentence is 0.35. Figure 4 The furthest vector distance of a single semantic sentence is 0.9, and the closest vector distance of a mixed conversation sentence is 0.2. Figure 5 In the example, the furthest vector distance of a single semantic sentence is 0.8, while the closest vector distance of a mixed conversation sentence is 0.4. Next, the mean distances of the furthest sentence vectors of the single semantic sentences with the same intent are calculated: A = (0.9 + 0.88 + 0.85 + 0.9 + 0.8) / 5 = 0.87, and the mean distances of the closest sentence vectors of the mixed conversation sentences with the same intent are calculated: B = (0.3 + 0.2 + 0.35 + 0.2 + 0.4) / 5 = 0.29. Finally, the average value of A and B is calculated: C = (0.87 + 0.29) / 2 = 0.58, and the preset distance threshold is determined to be 0.58.

[0101] Step S2054: The server determines the first conversation sentence from the first conversation text based on the third sentence.

[0102] After obtaining the third sentence, the server may determine the first conversation sentence in the first conversation text according to the third sentence.

[0103] In some embodiments, the first clustering result includes N first cluster sets, where N is an integer greater than 1; see Figure 6 , Figure 6This is a flowchart of determining a first conversation sentence from a first conversation text based on a third sentence provided by an embodiment of the present application. Figure 6 It is shown that in step S2054, the server determines the first conversation sentence from the first conversation text based on the third sentence, which can be achieved by the following steps S20541 to S20543:

[0104] In step S20541, the server performs the following processing on the third statement in any one of the N first cluster sets: the third statement and N-1 first cluster sets are combined into a third statement set, and any first cluster set is different from the corresponding N-1 first cluster sets.

[0105] As an example of step S20541, if the first sentence set has a total of 5 first cluster sets, the first first cluster set to which the third sentence A belongs is "a, b, c, d, e, f", where c is the first central sentence A and f is the third sentence A; the second first cluster set to which the third sentence B belongs is "1, 2, 3, 4, 5, 6", where 3 is the first central sentence B and 6 is the third sentence B; when determining the third sentence set, there are two cases:

[0106] The first case: delete the second sentence and the first central sentence A except the third sentence A in the first first cluster set to which the third sentence A belongs. At the same time, delete the second sentence and the first central sentence B except the third sentence B in the second first cluster set to which the third sentence B belongs; then, combine the third sentence A and the third sentence B and the remaining three sentences in the first cluster sets into a third sentence set, and proceed to the next step.

[0107] The second case: delete the second statement and the first center A except the third statement A in the first first cluster set to which the third statement A belongs, and combine the third statement A and the remaining four statements in the first cluster set into a third statement set; proceed to the next step; at the same time, delete the second statement and the first center sentence B except the third statement B in the second first cluster set to which the third statement B belongs, and combine the third statement B and the remaining four statements in the first cluster set into a third statement set, and proceed to the next step.

[0108] In step S20542, the server performs a clustering operation on the statements in the third statement set to obtain a second clustering result; a second cluster set in the second clustering result includes at least the third statement and the second central sentence.

[0109] In some embodiments, the server re-performs a clustering operation on the third set of statements to obtain a second clustering result, which includes at least the second central sentence and the third statement; the detailed process of the clustering operation has been described in step S2052, and this application will not repeat it here.

[0110] Step S20543: The server determines the first conversation sentence in the first conversation text based on the distance between the second central sentence and the third sentence.

[0111] Here, when the distance between the second central sentence and the third sentence is greater than a second preset distance threshold, the server determines the third sentence as the first conversation sentence in the first conversation text.

[0112] It should be noted that, in steps S2051 to S2054, the first clustering can preliminarily identify the special sentences (the third sentence mentioned above) in the first conversation text, but there may be misjudgments or omissions. The second clustering is used to verify the special sentences identified in the first time, which can effectively reduce the possibility of misjudgment and improve the accuracy of recognition of the first conversation sentences. Moreover, the two clusterings can enhance the robustness of recognition of the first conversation sentences. Even if the first clustering is affected by noise or sample distribution, the second clustering can further filter and confirm the special sentences, thereby further improving the recognition accuracy of the first conversation sentences and reducing the risk of misjudgment.

[0113] Step S206: The server determines, based on the attribute information of the first conversation object, a second conversation statement in the first conversation statement that belongs to the first conversation object.

[0114] The attribute information of the first conversation partner includes explicit characteristics (such as name, age, and gender), implicit characteristics (such as emotional state and conversation style), social relationships (such as friends and colleagues), and online behavior characteristics (such as behavior on social media, including likes, shares, and comments). By obtaining the attribute information of the first conversation partner, the accuracy of text processing can be improved.

[0115] In some embodiments, see Figure 7 , Figure 7 This is a flowchart of determining a second conversation statement based on attribute information of a first conversation object provided by an embodiment of the present application. Figure 7 In step S206, the server determines the second conversation statement belonging to the first conversation object in the first conversation statement according to the attribute information of the first conversation object, which can be achieved by the following steps S2061 to S2064:

[0116] Step S2061: The server divides the first conversation sentence into a set of sentence fragments.

[0117] A sentence fragment refers to a smaller unit segmented from the first conversation sentence. Each sentence fragment usually contains a relatively complete meaning or semantic unit.

[0118] In some embodiments, the first conversational sentence can be segmented in a variety of ways. Common segmentation methods include: punctuation-based segmentation, grammatical structure-based segmentation, semantics-based segmentation, and machine learning-based segmentation. Different segmentation methods can be selected based on different application scenarios. For example, to simply segment a long sentence into multiple clauses, punctuation-based segmentation can be selected, using punctuation as a dividing point to segment the sentence into multiple segments. When the sentence structure is more complex, grammatical structure-based segmentation can be selected. Syntactic analysis tools can be used to analyze the sentence's grammatical structure tree and segment the sentence according to basic structures such as subject, predicate, and object. For example, a complex sentence can be segmented into subordinate clauses and main clauses.

[0119] Step S2062: The server extracts features from the attribute information of the first session object to obtain an object feature vector.

[0120] An object feature vector represents features extracted from attribute information as a mathematical vector, facilitating subsequent calculations and comparisons. The server can extract the object feature vector using a vector extraction model, which can be either a machine learning model or a deep learning model.

[0121] In step S2063 , the server performs feature extraction on the sentence fragments in the sentence fragment set to obtain sentence fragment vectors.

[0122] Sentence fragment vectors represent features extracted from sentence fragments as mathematical vectors. Sentence fragment vectors can also be extracted using the vector acquisition model.

[0123] Step S2064: The server determines a second conversation sentence belonging to the first conversation object from the sentence fragment set based on the object feature vector and the sentence fragment vector.

[0124] The server analyzes the object feature vector and the sentence fragment vector to identify the second conversation sentence belonging to the first conversation object in the sentence fragment set.

[0125] In some embodiments, step S2064 can be implemented by the following method: first, performing vector splicing on the object feature vector and the sentence fragment vector to obtain a feature splicing vector; then, based on the feature splicing vector, determining the probability value that the sentence fragments in the sentence fragment set belong to the first conversation object; finally, based on the probability value, determining the second conversation sentence belonging to the first conversation object from the sentence fragment set.

[0126] Vector concatenation involves joining two or more vectors to form a new, longer vector. Vector concatenation can be simple dimension-wise addition or weighted concatenation. For example, if the sentence fragment vector is [a, b, c] and the object feature vector is [e, f], then the concatenated vector obtained by dimension-wise addition is [a, b, c, e, f].

[0127] In some embodiments, based on the feature splicing vector, determining the probability value that a sentence fragment in the sentence fragment set belongs to the first conversation object can be achieved by the following method: first, the server inputs the feature splicing vector into a trained classification model, which can be a neural network model. The features of the input feature splicing vector are gradually extracted and learned through the convolution layer and pooling layer in the classification model to generate representation features for the feature splicing vector. Then, the generated representation features are input into the fully connected layer in the classification model. The extracted representation features are converted into probability values ​​through the activation function in the fully connected layer. Finally, the probability values ​​of the sentence fragments in the sentence fragment set belonging to the first conversation object are output.

[0128] In some embodiments, determining the probability value of a sentence fragment in the sentence fragment set belonging to the first conversation object based on the feature splicing vector can also be achieved by the following method: the server can also input the feature splicing vector into a pre-trained large model (such as qwen2-72B), and guide the pre-trained large model to output the probability value of each sentence fragment in the sentence fragment set belonging to the first conversation object through a prompt word. For example, the prompt word is "You are a probability value determination model, which outputs the probability value of the sentence fragment belonging to the first conversation object based on the input feature splicing vector." After receiving the prompt word, the pre-trained large model can parse the feature splicing vector to obtain an object feature vector and a sentence fragment vector; then, calculate the similarity between the object feature vector and the sentence fragment vector, and output the probability value of the sentence fragment corresponding to the sentence fragment vector belonging to the first conversation object.

[0129] In some embodiments, determining the second conversation sentence belonging to the first conversation object from the sentence fragment set based on the probability value can be achieved by the following method: first, obtaining a preset probability threshold, then comparing the obtained probability value of the sentence fragment belonging to the first conversation object with the preset probability threshold; when the probability value is greater than the preset probability threshold, it indicates that the sentence fragment is the second conversation sentence belonging to the first conversation object.

[0130] In some embodiments, the server may also determine the relationship between the first conversation object and the other conversation objects based on the first conversation statement. This may be achieved by the following method: first, the server identifies a term of address in the first conversation statement; then, based on the identified term of address, the server determines the relationship between the first conversation object and the other conversation objects. For example, if the term of address "mother" is identified in the first conversation statement, the server may determine that the first conversation object and the other conversation objects are in a mother-child relationship.

[0131] Step S207: The server updates the first conversation text based on the second conversation statement to obtain a second conversation text belonging to the first conversation object.

[0132] The specific updating method is shown in step S103, which will not be described in detail in this application.

[0133] Step S208: The server determines the second conversation text as the text processing result.

[0134] Step S209: The server sends the text processing result to the terminal.

[0135] Step S210: The terminal displays the text processing result on the current interface.

[0136] After obtaining the text processing results, you can perform the next step based on the text processing results, such as intent recognition.

[0137] The text processing method provided in the embodiment of the present application can recognize and process voice input of different language types, thereby adapting to different application scenarios and meeting diverse user needs. It can also be expanded to more languages ​​and text formats, thereby improving the scalability and application scope of the system. In addition, all texts are unified into the same language or format, which facilitates subsequent processing and analysis, thereby improving the accuracy and efficiency of text processing. In addition, by identifying the first conversation sentence in the first conversation text, effective data support can be provided for subsequent text processing, thereby improving the efficiency and quality of text processing. In addition, by combining the attribute information of the first conversation object, the second conversation text belonging to the first conversation object can be more accurately identified in the first conversation sentence, thereby improving the recognition accuracy and reducing the probability of misjudgment, thereby efficiently and accurately removing the conversation text that does not belong to the first conversation object in the first conversation sentence.

[0138] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.

[0139] The conversation partners (the first conversation partner) may unconsciously express language differences due to factors such as their jobs and educational levels. This embodiment of the present application combines the feature information of the conversation partners (the attribute information of the first conversation partner) to infer multiple conversation partners in the first conversation text. The following describes the application of this embodiment of the present application in a bank debt collection scenario.

[0140] See also Figure 8 , Figure 8 This is a text processing framework diagram provided by an embodiment of the present application. Specifically, it can be implemented by following steps S401 to S407:

[0141] Step S401: Acquire all historical conversation text data of the debt collection scenario.

[0142] The historical conversation text data (corresponding to the first conversation text mentioned above) refers to the conversation text data generated before the current time, and the conversation text data can be obtained by converting voice data.

[0143] Historical conversation text data can be in the following forms:

[0144] Text 1: Hello, oh, wait, phone call, hello, this is me.

[0145] Text 2: I am myself.

[0146] Text 3: I have settled.

[0147] Text 4: Wait.

[0148] Text 5: Phone.

[0149] Step S402: Input the historical conversation text data into a semantic confusion recognition model to determine whether there is a semantic confusion sentence (the first conversation sentence mentioned above).

[0150] The semantic confusion recognition model actually performs preliminary semantic judgment on historical conversation text data by clustering. Step S402 can be implemented by the following method:

[0151] First, the BERT+K-means clustering algorithm is used to cluster the historical conversation text data of the same conversation subject (the first conversation subject mentioned above). Specifically, BERT is used to obtain sentence vectors for the historical conversation text data. Clustering is then performed based on vector distance to obtain a first clustering result. The first clustering result includes multiple first sentence cluster sets. Each first cluster set contains the conversation text (corresponding to the second sentence mentioned above), the cluster number, the central sentence (corresponding to the first central sentence mentioned above), and the distance between the conversation text and the central sentence. Then, conversation texts whose distance from the central sentence exceeds a first preset distance threshold are selected; the central sentence is denoted as Center1, the cluster is denoted as Cluster1 (the first cluster set mentioned above), and the selected special text (the third sentence mentioned above) is denoted as Sent1. For example, if the distance between Text 1 and the central sentence Text 2 exceeds the first preset distance threshold, Text 1 is determined to be a special text. Next, the text in cluster Cluster1 (step 2) except for the special text Sent1 is deleted. Re-clustering is performed using the special text Sent1 and all text data in the remaining clusters (the third sentence set mentioned above) to obtain a second clustering result. Next, the cluster (the aforementioned second cluster set) to which the special text sent1 belongs in the second clustering result is determined. For example, the special text sent1 belongs to cluster cluster2 in the second clustering result, and the central sentence of cluster2 is text 4. Finally, it is determined whether the distance between the special text sent1 and the central sentence of the new cluster is greater than a second preset distance threshold. If so, the special text sent1 is determined to be a semantically mixed sentence.

[0152] Step S403: Acquire characteristic information of the conversation object to which the historical conversation text data belongs (attribute information of the first conversation object mentioned above).

[0153] Feature information includes the conversation partner's age, gender, family members, occupation, etc.

[0154] Step S404: construct different voice-over combinations according to the characteristic information of the conversation object.

[0155] For example: The voice-over combination can be mother + myself, child + myself.

[0156] Step S405 , based on the feature information of the conversation object and the voice-over combination, determining the conversation text that does not belong to the conversation object in the semantically mixed sentence, and the relationship between the conversation object and other conversation objects.

[0157] Step S406: retain the conversation text (the second conversation statement) belonging to the conversation object, and delete the conversation text that does not belong to the conversation object, thereby stripping the historical conversation text data.

[0158] Step S407: The stripped data (the second conversation text) is transferred into the intention recognition training data set.

[0159] In the embodiment of the present application, a social dialect recognition model is formed from the encoding at the vocabulary, sentence and other levels to achieve the automation of text stripping; by calling the social dialect recognition model outside the intention recognition model, the conversation object and other conversation objects to which the historical conversation text data belongs, that is, the corresponding conversation text, are determined, the stripping of semantically mixed sentences is completed, and the intention recognition is assisted, providing technical support for improving the accuracy of intention recognition.

[0160] It is understandable that in the embodiments of the present application, if data related to user information or corporate information is involved, when the embodiments of the present application are applied to specific products or technologies, it is necessary to obtain user permission or consent, or to blur this information to eliminate the correspondence between this information and the user; and the relevant data collection and processing should be strictly in accordance with the requirements of relevant laws and regulations when applied in examples, and the informed consent or separate consent of the personal information subject should be obtained, and subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.

[0161] The following continues to describe the exemplary structure of the text processing device provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 9 As shown, the text processing device 254 includes: a first determination module 2541, which is used to perform semantic recognition on a first conversation text of a first conversation object and determine a first conversation sentence in the first conversation text, where the first conversation sentence includes conversation sentences of multiple conversation objects; a second determination module 2542, which is used to determine a second conversation sentence in the first conversation sentence belonging to the first conversation object based on attribute information of the first conversation object; and an updating module 2543, which is used to update the first conversation text based on the second conversation sentence to obtain a second conversation text belonging to the first conversation object.

[0162] In an embodiment of the present application, the first determination module is further used to: perform sentence division on the first conversation text to obtain a first sentence set; perform a clustering operation on the sentences in the first sentence set to obtain a first clustering result; each first cluster set of the first clustering result includes a first central sentence and K second sentences; K is a positive integer; based on the distance between the second sentence and the first central sentence, determine a third sentence in each first cluster set from the K second sentences; the distance between the third sentence and the first central sentence is greater than a first preset distance threshold; based on the third sentence, determine the first conversation sentence from the first conversation text.

[0163] In an embodiment of the present application, the first clustering result includes N first cluster sets, where N is an integer greater than 1; the first determination module is further used to: perform the following processing on the third sentence in any one of the N first cluster sets: combine the third sentence and N-1 first cluster sets to form a third sentence set, and any first cluster set is different from the corresponding N-1 first cluster sets; perform a clustering operation on the sentences in the third sentence set to obtain a second clustering result; a second cluster set in the second clustering result includes at least the third sentence and the second central sentence; and determine the first conversation sentence in the first conversation text based on the distance between the second central sentence and the third sentence.

[0164] In an embodiment of the present application, the first determination module is further configured to: determine the third sentence as the first conversation sentence in the first conversation text when the distance between the second central sentence and the third sentence is greater than a second preset distance threshold.

[0165] In an embodiment of the present application, the second determination module is further configured to: divide the first conversation sentence to obtain a set of sentence fragments; perform feature extraction on attribute information of the first conversation object to obtain an object feature vector; perform feature extraction on sentence fragments in the set of sentence fragments to obtain a sentence fragment vector; and determine, based on the object feature vector and the sentence fragment vector, a second conversation sentence belonging to the first conversation object from the set of sentence fragments.

[0166] In an embodiment of the present application, the second determination module is further used to: perform vector splicing on the object feature vector and the sentence fragment vector to obtain a feature splicing vector; determine, based on the feature splicing vector, a probability value that a sentence fragment in the sentence fragment set belongs to the first conversation object; and determine, based on the probability value, a second conversation sentence belonging to the first conversation object from the sentence fragment set.

[0167] In an embodiment of the present application, the device also includes a conversion module, which is used to: recognize the conversational speech of the first conversation object to obtain the language type of the conversational speech; when the language type is the first type, convert the conversational speech into the first conversational text through a first model; when the language type is the second type, convert the conversational speech into a third conversational text under the second type through the first model; and perform text language category conversion processing on the third conversational text to obtain the first conversational text of the first type.

[0168] In an embodiment of the present application, the conversion module is further used to: replace the second type of vocabulary in the third conversation text with the corresponding first type of vocabulary according to a preset mapping relationship between the second type of vocabulary and the first type of vocabulary to obtain a conversation conversion text; standardize the sentence structure in the conversation conversion text according to a predefined sentence conversion rule to obtain a standardized conversation text; and determine the standardized conversation text as the first conversation text.

[0169] It should be noted that the description of the device embodiment of the present application is similar to the description of the method embodiment described above, and has similar beneficial effects as the method embodiment, so it will not be repeated. For technical details not disclosed in the device embodiment, please refer to the description of the method embodiment of the present application for understanding.

[0170] The embodiment of the present application provides an electronic device to implement the above-mentioned text processing method. Figure 10 Schematic diagram of the structure of the electronic device provided in the embodiment of the present application. Figure 10 As shown, the electronic device 1200 includes at least a processor 1201 , a communication interface 1202 and a storage medium 1203 configured to store executable instructions, wherein the processor 1201 generally controls the overall operation of the electronic device 1200 .

[0171] The communication interface 1202 enables the electronic device to communicate with other terminals or servers through a network.

[0172] The storage medium 1203 is configured to store instructions and applications executable by the processor 1201, and can also cache data to be processed or processed by the processor 1201 and each module in the electronic device 1200, which can be implemented through flash memory (FLASH) or random access memory (RAM).

[0173] An embodiment of the present application provides a computer program product, which includes a computer program or executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the text processing method described in the embodiment of the present application.

[0174] The embodiment of the present application provides a computer-readable storage medium storing executable instructions, wherein a computer program or computer-executable instructions are stored. When the computer-executable instructions or computer program are executed by a processor, the processor will be caused to execute the text processing method provided by the embodiment of the present application, for example, Figure 2 The method shown.

[0175] In some embodiments, the storage medium can be a computer-readable storage medium, such as a ferroelectric random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPR OM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); it can also be various devices including one or any combination of the above memories.

[0176] In some embodiments, executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0177] As an example, the executable instructions may, but need not necessarily, correspond to a file in a file system, may be stored as part of a file storing other programs or data, for example, in one or more scripts in a Hypertext Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions). As an example, the executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located in one location, or on multiple electronic devices distributed in multiple locations and interconnected by a communication network.

[0178] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. A text processing method, characterized in that: The method comprises: Performing semantic recognition on a first conversation text of a first conversation object to determine a first conversation sentence in the first conversation text, where the first conversation sentence includes conversation sentences of a plurality of conversation objects; determining, according to the attribute information of the first conversation object, a second conversation statement in the first conversation statement that belongs to the first conversation object; The first conversation text is updated based on the second conversation statement to obtain a second conversation text belonging to the first conversation object.

2. The method according to claim 1, characterized in that The performing semantic recognition on the first conversation text of the first conversation object to determine the first conversation sentence in the first conversation text includes: Dividing the first conversation text into sentences to obtain a first sentence set; Performing a clustering operation on the sentences in the first sentence set to obtain a first clustering result; each first cluster set of the first clustering result includes a first central sentence and K second sentences; K is a positive integer; Based on the distance between the second sentence and the first central sentence, determining a third sentence in each first cluster set from the K second sentences; the distance between the third sentence and the first central sentence is greater than a first preset distance threshold; Based on the third sentence, the first conversation sentence is determined from the first conversation text.

3. The method according to claim 2, characterized in that The first clustering result includes N first cluster sets, where N is an integer greater than 1; and determining the first conversation sentence from the first conversation text based on the third sentence includes: Performing the following processing on the third statement in any one of the N first cluster sets: combining the third statement with the N-1 first cluster sets to form a third statement set, where any one of the first cluster sets is different from the corresponding N-1 first cluster sets; performing a clustering operation on the sentences in the third sentence set to obtain a second clustering result; wherein any second cluster set in the second clustering result includes at least the third sentence and the second central sentence; The first conversation sentence in the first conversation text is determined based on the distance between the second central sentence and the third sentence.

4. The method according to claim 3, characterized in that The determining, based on the distance between the second central sentence and the third sentence, the first conversation sentence in the first conversation text includes: When the distance between the second central sentence and the third sentence is greater than a second preset distance threshold, the third sentence is determined to be the first conversation sentence in the first conversation text.

5. The method according to claim 1, wherein The determining, based on the attribute information of the first conversation object, a second conversation statement in the first conversation statement that belongs to the first conversation object includes: Dividing the first conversation sentence to obtain a set of sentence fragments; Extracting features from the attribute information of the first conversation object to obtain an object feature vector; Extracting features from the sentence fragments in the sentence fragment set to obtain sentence fragment vectors; Based on the object feature vector and the sentence fragment vector, a second conversation sentence belonging to the first conversation object is determined from the sentence fragment set.

6. The method according to claim 5, characterized in that The determining, based on the object feature vector and the sentence fragment vector, a second conversation sentence belonging to the first conversation object from the sentence fragment set includes: Performing vector concatenation on the object feature vector and the sentence fragment vector to obtain a feature concatenation vector; Determining, based on the feature concatenation vector, a probability value of a sentence fragment in the sentence fragment set belonging to the first conversation object; Based on the probability value, a second conversation sentence belonging to the first conversation object is determined from the sentence fragment set.

7. The method according to claim 1, characterized in that Before performing semantic recognition on the first conversation text of the first conversation object, the method further includes: Recognizing the conversational speech of the first conversation partner to obtain a language type of the conversational speech; In a case where the language type is the first type, converting the conversation speech into the first conversation text by using a first model; When the language type is the second type, the conversation speech is converted into a third conversation text of the second type through the first model; and the third conversation text is subjected to text language category conversion processing to obtain the first conversation text of the first type.

8. An electronic device, characterized in that: include: a memory for storing executable instructions; The processor is configured to implement the text processing method according to any one of claims 1 to 7 when executing the executable instructions or computer program stored in the memory.

9. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the text processing method according to any one of claims 1 to 7 is implemented.

10. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the text processing method according to any one of claims 1 to 7 is implemented.