Method for unsupervised analysis of a textual data set related to the execution of business processes

EP4639397A1Pending Publication Date: 2025-10-29ORANGE SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2023828206
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-19
Filing Date
2023-12-18
Publication Date
2025-10-29

AI Technical Summary

Technical Problem

Current methods for analyzing unstructured textual data from business applications, such as customer feedback and complaints, require substantial human investment and are limited in granularity, failing to automatically discover new categories or provide detailed reasons for user sentiments.

Method used

A method that identifies recurring word combinations, groups synonymous words into semantic entities, determines typical actions, and classifies them into sub-themes and themes, allowing for unsupervised analysis and finer granularity in data extraction, while accounting for spelling errors and derivation relationships.

Benefits of technology

Enables automatic, unsupervised analysis of textual data to extract relevant business knowledge at a fine level of granularity, reducing human intervention and providing structured information on user sentiments and their reasons, facilitating better data exploitation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 1.1
    Figure 1.1
Patent Text Reader

Abstract

The invention relates to a method for analysing a textual data set referred to as a verbatim set (EV), each verbatim comprising an ordered sequence of words. Such a method comprises: - a step (12) of determining recurring combinations of words in the verbatims and of identifying common patterns within the recurrent combinations of words; - a step (13) of identifying synonymous relationships between words having the same syntactic function within the common patterns and of grouping the synonymous words having the same syntactic function within the same data structures, which data structures are referred to as semantic entities; - a step (14) of determining type actions according to the semantic entities, a type action being representative of similar contexts of operations identifiable in the verbatims.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DESCRIPTION

[0002] TITLE: Unsupervised analysis method for a set of textual data linked to the execution of business processes

[0003] Technical field

[0004] The invention lies in the field of automatic analysis of unstructured textual data. More particularly, the invention relates, for example, to the analysis of textual data from business application logs.

[0005] Prior art

[0006] In many fields, and more particularly in the context of business management, business applications are often used to collect textual data freely written by a user in natural language, for example within an input area of ​​a form made available via a computer tool. Examples of such textual data include verbatim statements from various interlocutors (e.g. customers, employees, etc.) collected via customer relationship management applications, comments from incident escalation tickets, exchanges tracked within a customer complaint management system, etc.

[0007] The unstructured nature of this data, however, makes it particularly difficult to exploit, at least in a completely automatic manner. Therefore, laborious and time-consuming manual operations are often necessary to extract useful information from this textual data, such as a user's feelings (e.g., their satisfaction or dissatisfaction) or the purpose and reasons for this feeling.

[0008] Text data mining techniques have been developed to facilitate these operations, but these existing solutions, which can be divided into two broad categories - solutions based on a supervised learning approach on the one hand, and solutions based on an unsupervised learning approach on the other hand - generally have several drawbacks, as detailed below.

[0009] A main drawback of supervised learning-based approaches is that they require a particularly large human investment to build a labeled learning base large enough to allow obtaining reliable results. More precisely, each sample of such a learning base must be manually annotated, which is extremely time-consuming. Another important drawback of these supervised learning-based solutions lies in the fact that the classification carried out is carried out on the basis of predefined categories, i.e. previously identified and selected by a human operator. In other words, these solutions do not allow the automatic discovery of new categories (e.g. new feelings, new objects of a feeling, etc.).) which would appear over time and which would not have been manually identified from the start in the learning base used.

[0010] Current approaches based on unsupervised learning only partially overcome some of the aforementioned drawbacks. While they require a lower overall human investment than that required for the implementation of solutions based on supervised learning, this investment is still relatively substantial, since the construction of a large learning database is still required (even if it is not necessary in this case to annotate it manually).

[0011] Furthermore, regardless of the approach considered, we also note that the level of granularity of the results generated by existing solutions is often relatively coarse. For example, many of these existing solutions are limited to automatically determining a user feeling or sentiment, without necessarily providing information on the reasons that could explain this feeling. Also, significant human intervention often remains necessary even after the implementation of these automated solutions, in order to break down the results obtained into finer levels of granularity and thus allow a more interesting and relevant exploitation of the available textual data.

[0012] There is therefore a need for a solution to improve the analysis of unstructured text data collected via applications or business processes, in particular by enabling the implementation of more detailed automated analysis of this data.

[0013] Summary of the invention

[0014] The present technique makes it possible to propose a solution aimed at remedying certain drawbacks of the prior art. According to a first aspect, the present technique relates in fact to a method for analyzing a set of textual data, called a set of verbatims, each verbatim comprising an ordered sequence of words, this method comprising: a step of determining recurring word combinations in said verbatims, and of identifying common patterns within said recurring word combinations; a step of identifying synonymous relationships between words of the same syntactic function within said common patterns, and of grouping said synonymous words of the same syntactic function within the same data structures, called semantic entities; a step of determining typical actions as a function of said semantic entities, a typical action being representative of similar contexts of operations identifiable in said verbatims.

[0015] In this way, the present technique makes it possible to identify, from a set of verbatims, typical actions representative of various operating contexts detected in the verbatims. Thus, the present technique makes it possible to draw up, at a fine degree of granularity and in an unsupervised manner, an overview of the different operating contexts identifiable in a set of verbatims.

[0016] In a particular embodiment, said grouping into semantic entities further comprises taking into account derivation relationships and / or spelling corrections between words of said common patterns.

[0017] In this way, the present technique offers a certain flexibility when analyzing textual data, taking into account potential reformulations or spelling errors in the terms used in the verbatims.

[0018] In a particular embodiment, said semantic entities comprise semantic entities of the verb, noun, or adjective type, and a typical action comprises a semantic entity of the verb type and at least one semantic entity of the noun or adjective type.

[0019] In this way, the semantic entities are categorized, allowing a fine analysis of the data identified in the verbatims, depending on whether they represent an object (semantic entity of noun type), an action (semantic entity of verb type), or a description (semantic entity of adjective type).

[0020] According to a particular characteristic of this embodiment, said method comprises the association of said typical actions with said verbatims, a typical action being associated with a verbatim when a combination of words present in said verbatim is associated with a common pattern comprising at least one word associated with the verb-type semantic entity of said typical action and at least one word associated with said at least one noun or adjective-type semantic entity of said typical action.

[0021] In this way, the typical actions are linked to the verbatims in which they appear.

[0022] In a particular embodiment, said determination of typical actions comprises the identification of conditional relationships between said semantic entities, based on a probability of coexistence of words of said semantic entities within the same verbatims.

[0023] In this way, complementary typical actions are identified on the basis of dependency criteria identified in certain verbatims

[0024] According to a particular characteristic of this embodiment, said determination of typical actions comprises the construction of at least one graph as a function of said conditional relationships, and the use of said at least one graph to group said typical actions into sub-themes, and said sub-themes into themes.

[0025] In this way, an automatic classification of typical actions into sub-themes and themes is implemented, allowing the obtaining of an overview of the different contexts of operations identifiable in a set of verbatims at different degrees of granularity, including at least a fine degree of granularity (i.e. categorization into typical actions) and more general degrees of granularity (categorization into sub-themes and themes).

[0026] In a particular embodiment, said method further comprises a step of segmenting said verbatims into text segments, and of determining a polarity associated with each of said text segments.

[0027] In this way, the present technique also makes it possible to associate different parts of the verbatims with categories of feeling expressed there by users, for example a negative feeling (i.e. the expression of dissatisfaction), a neutral feeling (i.e. neither satisfied nor dissatisfied), or a positive feeling (i.e. the expression of satisfaction).

[0028] According to a particular characteristic of this embodiment, said method comprises a step of weighting said typical actions, a typical action being weighted according to the polarities associated with the text segments of the verbatims with which said typical action is associated.

[0029] In this way, the present technique not only allows to establish a link between polarities and typical actions, that is to say between a user feeling and a reason for this feeling, but also to identify trends by taking into account the number of textual segments within which a particular feeling is associated with a particular reason.

[0030] According to another aspect, the present technique also relates to an electronic device for analyzing a set of textual data, called a set of verbatims, each verbatim comprising an ordered sequence of words, said device comprising: means for determining recurring word combinations in said verbatims, and for identifying common patterns within said recurring word combinations; means for identifying synonymous relationships between words of the same syntactic function within said common patterns, and for grouping said synonymous words of the same syntactic function within the same data structures, called semantic entities; means for determining typical actions, as a function of conditional relationships identified between said semantic entities within said verbatims, a typical action being representative of similar contexts of operations identifiable in said verbatims.

[0031] The means of said electronic device may be adapted to the implementation of any of the embodiments of the method of the present application.

[0032] According to another aspect, the proposed technique also relates to a computer program product downloadable from a communications network and / or stored on a computer-readable medium and / or executable by a microprocessor, comprising program code instructions for executing an analysis method as described above, when executed on a computer.

[0033] The proposed technique also relates to a computer-readable recording medium on which is recorded a computer program comprising program code instructions for executing the steps of the method as described above, in any of its embodiments. Such a recording medium may be any entity or device capable of storing the program. For example, the medium may comprise a storage means, such as a ROM, for example a CD ROM or a microelectronic circuit ROM, or a magnetic recording means, for example a USB key or a hard disk.

[0034] On the other hand, such a recording medium may be a transmissible medium such as an electrical or optical signal, which may be conveyed via an electrical or optical cable, by radio or by other means, so that the computer program contained therein is remotely executable. The program according to the invention may in particular be downloaded over a network, for example the Internet.

[0035] The different embodiments mentioned above can be combined with each other to implement the invention.

[0036] Figures

[0037] Other characteristics and advantages of the invention will appear more clearly on reading the following description of a preferred embodiment, given as a simple illustrative and non-limiting example, and the appended drawings, among which:

[0038] [Fig 1] illustrates the main steps of a textual data analysis process, in a particular embodiment of the proposed technique;

[0039] [Fig 2] presents an example of an extract of a directed graph obtained in the context of the analysis of a set of verbatims, in a particular embodiment of the proposed technique; [Fig 3] describes a simplified architecture of an electronic device for the implementation of the proposed technique, in a particular embodiment.

[0040] Detailed description of the invention

[0041] This application makes it possible to remedy some of the aforementioned drawbacks.

[0042] More specifically, a technique is proposed for analyzing, in an automatic and unsupervised manner, a corpus of unstructured textual data, in order to extract relevant information, for example business knowledge. Although this field of use is given for illustrative and non-limiting purposes, the present technique is indeed advantageous for facilitating the exploitation of data from logs generated at the level of business applications now widely deployed within companies (e.g. customer relationship management applications, incident management, etc.). The business knowledge that the present technique makes it possible to extract includes, for example, the identification of problems in the execution of certain business processes, the determination of the returns (or "feedback" in English) and / or the feelings of the actors involved in these business processes (e.g.reasons for satisfaction or dissatisfaction), or the determination of the recurring nature of certain business activities. The proposed technique offers a better compromise than existing solutions, in that it allows data extraction at a relatively fine level of granularity while limiting the human investment (i.e. the number of non-automated manual operations) necessary to achieve this result.

[0043] As detailed below, in a particular embodiment, the proposed technique also makes it possible to automatically structure the extracted information, by organizing it into automatically identified sub-themes and themes. In other words, the present technique is not limited to raw data extraction, but also aims to provide automatic structuring of relevant data according to different levels of granularity, i.e. at different levels of generalization, thus making their exploitation easier.

[0044] According to a first aspect, the present technique relates to a method for analyzing a set - or corpus - of textual data. More particularly, in the context of the present technique, these textual data are designated by the generic term "verbatims", a verbatim corresponding to a set of words entered by a user within, for example, a business application, with regard to one or more given objects (for example, a service, an incident, a service, etc.). A verbatim can thus be likened to a coherent set of words, carrying meaning and based on structures specific to a language, allowing, for example, a user to express a feeling (negative, neutral or positive), but also possibly the object and the reasons for this feeling. Such textual data are for example stored in the form of text files or within dedicated fields of one or more databases.

[0045] We now present, in relation to Figure 1, the general principle of a method for analyzing a set of verbatims, in a particular embodiment of the proposed technique. Such a method is implemented by an electronic device having access to at least one data source comprising a set EV of verbatims (VI, ..., Vn) to be processed. In a step 12, a determination of recurring word combinations within the verbatims is implemented, then common patterns are identified within these word combinations. According to a particular characteristic, a common pattern groups together, for example, the semantically equivalent word combinations among the determined word combinations.In other words, step 12 aims to discover word combinations that are recurrent in all the verbatims, and to group together the word combinations that are used to express the same general idea, by associating them with the same common motif.

[0046] Assuming, for example, that we identify in the verbatim a recurrence of the following word combinations: "fix a problem", "solve my problem", "solve this problem", "solve the problem", "solve a problem", the implementation of step 12 makes it possible to highlight that all these expressions form various semantic representations of the same idea - in this case the resolution of a problem - by associating them with the same common motif.

[0047] More specifically, two concepts (i.e. two particular groups of words) are identifiable in the previous example: a first concept linked to the action of finding a solution, characterized by the words "restore", "resolve", and "troubleshoot", associated with the syntactic function "verb"; a second concept linked to the notion of difficulty, characterized by the words "problem", and "concern", associated with the syntactic function "noun".

[0048] All the preceding combinations can thus be grouped under the same common label, i.e. the same common pattern MC grouping together a set of concepts, defined for example by a data structure which can take the following non-limiting form:

[0049] MC: {({'fix', 'troubleshoot', 'solve'}, 'verb'), ({'problem', 'concern'}, 'noun')}.

[0050] In a particular embodiment, step 12 is implemented by performing a succession of sub-steps comprising: generating, for each verbatim of the set of verbatims, possible combinations of words, called candidate combinations, as a function of a predetermined distance threshold and a length threshold; calculating the number of occurrences of each candidate combination in the set of verbatims, and filtering the candidate combinations as a function of their number of occurrences, such filtering being able, for example, to consist of eliminating the candidate combinations which do not appear sufficiently in the set of verbatims (i.e. whose number of occurrences is less than a predetermined threshold); merging the close candidate combinations (on the basis, for example, of prior operations of lemmatization of the words of the candidate combinations);the determination of the set of synonymous words for each word of the candidate combinations, by means of a synonym dictionary (taking for example the form of a lexical database) accessible from the electronic device; the use of the determined synonymous words to identify patterns common to candidate combinations: more particularly, two candidate combinations are considered similar and grouped within the same group associated with the same common pattern when these two candidate combinations have the same number of words having the same syntactic functions, and when for each word ml of one of the two candidate combinations there exists a word m2 in the other of the two candidate combinations such that the words ml and m2 are synonymous.;

[0051] Thus, several groups of candidate combinations, each associated with a common pattern, are obtained at the end of step 12.

[0052] In a step 13, data structures, called semantic entities, are constructed on the basis of the common patterns identified in step 12. By semantic entity, we mean more particularly a data structure comprising a set of synonymous words with the same syntactic function. Thus, step 13 comprises the identification, by means of a dictionary of synonyms, of synonymous relationships between words with the same syntactic function (verbs, nouns or adjectives) within the common patterns, and the grouping of synonymous words with the same syntactic function by semantic entity.

[0053] According to a particular characteristic, the semantic entities include: semantic entities of the “verb” type, also called action entities, grouping the synonymous verbs identified in the set of common patterns; semantic entities of the “adjective” type, also called description entities, grouping the synonymous adjectives identified in the set of common patterns; semantic entities of the “noun” type, also called nominal entities, grouping the synonymous nouns identified in the set of common patterns.

[0054] In a particular embodiment, the construction of the semantic entities comprises, in addition to taking into account synonymy relations, taking into account derivation relations and spelling correction relations between the words present within the common patterns.

[0055] More specifically, taking into account derivation relationships makes it possible to bring together derived words when implementing step 13. For example, if a common pattern includes the word “resolve”, the derived word “resolution” is also considered for the creation of semantic entities even if it has not been identified as such within the common patterns obtained at the end of step 12.

[0056] Similarly, taking into account spelling correction relationships makes it possible to bring an incorrectly spelled word closer to its correctly spelled equivalent when implementing step 13. For example, spelling approximations linked to the omission of accents or the presence of incorrect accents, or the substitution of certain characters associated with the same sound (eg ' / ' and 'c' and 'q', 'x' and 'c', etc.), are also considered for the creation of semantic entities: thus the words 'problem' and 'problem' are for example considered equivalent, and grouped in the same semantic entity.

[0057] Below, for purely illustrative purposes, are two examples of semantic entities that can be obtained in step 13.

[0058] In a first example, we assume for example that step 12 of identifying common patterns has allowed the highlighting, among others, of the following concepts (or groups of words): {'affaire_question', 'question', 'difficulté_problème', 'interrogation_question', 'problème_question', 'souci', 'probleme_souci', 'problème', 'interrogation'}. On the basis of the relationships previously presented (synonymy, derivation, spelling correction), step 13 allows the construction of a semantic entity of type "name" EN in which the following different names are grouped, because identified as being synonymous: EN: 'difficulté_interrogation_problème_probleme_ souci_question_affaire'

[0059] In a second example, we assume, for example, that step 12 of identifying common patterns has enabled the highlighting, among others, of the following concepts (or groups of words): {'agréable_aimable', 'sympathique', 'agréable_aimable_sympathique', 'gentil', 'agréable', 'aimable_sympathique', 'avenant_sympathique', 'aimable_avenant_cordial_gentH', 'bienveillant_sympathique'}. On the basis of the relationships previously presented (synonymy, derivation, spelling correction), step 13 allows the construction of a semantic entity of the ED "adjective" type in which the following different adjectives are grouped together, because they are identified as being synonymous:

[0060] ED: 'pleasant_friendly_benevolent_cordial_kind_friendly_sympathetic'

[0061] Thus, several semantic entities of different types are obtained at the end of step 13.

[0062] In a step 14, the semantic entities constructed in step 13 are used to determine typical actions representative of similar contexts of operations (or use) identifiable in the EV set of analyzed verbatims.

[0063] According to a particular characteristic, a typical action comprises a semantic entity of verb type and at least one semantic entity of noun or adjective type.

[0064] More specifically: the verb-type semantic entity (or action entity) defines the type of operation associated with the action (e.g., an installation, a modification, a cancellation, etc.); the noun-type semantic entity (or nominal entity) defines the object that undergoes the effect of the type of operation carried out (e.g., fiber, an appointment, etc.); the adjective-type semantic entity (or description entity) defines the way in which the type of operation was carried out (e.g., late, incorrect, etc.).

[0065] Several grouping methods can be used to determine typical actions.

[0066] First, in a particular embodiment, a first grouping mode relates to the common patterns identified in step 12.

[0067] More specifically, the aim here is to associate with a given action type (i.e. to group together) the common patterns which simultaneously include: at least one word (i.e. a concept) belonging to or having a derivative belonging to the semantic entity of the verb type of the action type considered; at least one other word (i.e. another concept) belonging to or having a derivative belonging to the semantic entity of the noun or adjective type of the action type considered.

[0068] Thus, for example, according to this first grouping method: the common patterns 'install fibre' and 'installation fibre' are associated with the same action type, because the word 'installation' is part of the derivatives of the verb-type semantic entity (i.e. of the action entity) including the verb 'install'; the common patterns 'answer interrogation' and 'answer problem' are associated with the same action type, because on the one hand the word 'answer' is part of the derivatives of the verb-type semantic entity (i.e. of the action entity) including the verb 'answer', and on the other hand the words 'interrogation' and 'problem' belong to the same noun-type semantic entity (i.e. to the same nominal entity).

[0069] Secondly, in another particular embodiment complementary or alternative to the previous one, a second grouping mode is operated, on the basis of conditional relationships between semantic entities, identified within the verbatims. According to a particular characteristic, these conditional relationships between semantic entities are identified according to a probability of coexistence of words of said semantic entities within the same verbatims.

[0070] In this context, two types of conditional relations are considered for example: so-called global conditional relations, characterizing the identification of dependency relations between semantic entities of the noun or adjective type within the same verbatims; so-called contextual conditional relations, characterizing the identification of dependency relations between semantic entities of the noun or adjective type within the same verbatims, in the presence of a particular action semantic entity.

[0071] According to a particular characteristic, a conditional relationship between two semantic entities of the noun or adjective type is considered to be confirmed when the probability of coexistence of words of said semantic entities within the same verbatims is greater than a certain predetermined threshold, for one or the other of the two types of conditional relationships previously presented. On the basis of the identified conditional relationships, new action types are determined, a typical action comprising on the one hand the set of semantic entities of the noun or adjective type which are linked by global or contextual conditional relationships associated with the presence of the same semantic entity of the verb type, and on the other hand the semantic entity of the verb type in question. An association of certain patterns common to these new action types is also implemented, a common pattern being associated with a given action type when it comprises: at least one word (i.e.a concept) belonging to or having a derivative belonging to the verb-type semantic entity of the action type considered; at least one other word (i.e. another concept) belonging to or having a derivative belonging to a noun or adjective-type semantic entity of the action type considered.

[0072] In a particular embodiment, the determination of the typical actions comprises the construction of at least one graph based on the identified conditional relationships, and the use of this graph to group the typical actions into sub-themes, and the sub-themes into themes.

[0073] An example of an extract from such a graph is presented in relation to Figure 2, in a particular embodiment of the proposed technique. This directed graph is for example constructed from the previously identified global conditional relationships. More particularly, in this example given for purely illustrative purposes, each node of the graph corresponds to a semantic entity of the name type, and each directed link between two nodes corresponds to a global conditional relationship identified between the corresponding nodes. Each link is also associated with a weight, corresponding to the number of times the linked nodes have together formed contextual conditional relationships in the context of different verb-type semantic entities (i.e. action entities).For example, if in all the analyzed verbatims the nominal semantic entities "fiber" and "sheath" are always associated together in the presence of one or the other of the action entities associated with the verbs "install" and "verify", the weight "2" is associated with the link between these two nominal semantic entities, representative of the fact that two distinct contexts of operation (or use) have been identified in the verbatims in connection with these nominal semantic entities (i.e. a context in which these objects are mentioned in the context of an installation operation, and another in which these objects are mentioned in the context of a verification operation).

[0074] A graph constructed in this way is interesting in that it can be used to automatically determine possible groupings between semantic entities of the same type (and more particularly of the noun or adjective type), at different levels of granularity.

[0075] Thus, on the basis of predetermined rules, such a graph makes it possible, for example, to identify sub-themes, a sub-theme being representative of a set of semantic entities of the noun or adjective type which coexist in the same verbatims (without necessarily being the object of the same action entities), or in other words, which share the same context of operation.

[0076] According to an example of construction, a sub-theme groups together for example all the child nodes which descend from exactly the same parent nodes.

[0077] For example, applied to the graph extract illustrated in Figure 2, such a construction rule allows the automatic obtaining (i.e. without user intervention) of four sub-themes: the sub-theme ST11 comprising the nominal semantic entities “electrician” and “hole” having as parent node the nominal semantic entity “cable_cable_line”; the sub-theme ST12 comprising the nominal semantic entities “duct”, “sheath” and “adsl” having as parent nodes the nominal semantic entities “fiber” and “cable_cable_line”; the sub-theme ST13 comprising the nominal semantic entities “welding” and “garage” having as parent node the nominal semantic entity “.fiber”; the sub-theme ST21 comprising the nominal semantic entities “wifi”, “laptop”, “computer” and “decoder” having as parent node the nominal semantic entity “appliance_telephone_tv_television”.

[0078] Optionally, an additional rule excluding from a sub-theme the child nodes associated with a link of zero weight is also applied (such a rule implemented in the context of the previous example then leading for example to the exclusion of the nominal semantic entity "welding" from the sub-theme ST13, the weight of the link between the semantic entity "welding" and its parent the semantic entity "fiber" being zero). According to a particular characteristic, the sub-themes thus identified are in turn grouped into themes, on the basis of a predetermined rule.

[0079] According to an example of construction, a theme is for example defined by the union of sub-themes sharing common parent nodes.

[0080] For example, applied to the graph extract illustrated in Figure 2, such a construction rule allows the automatic obtaining (i.e. without user intervention) of two themes: the theme T1 grouping the sub-themes ST11, ST12 and ST13 which have as parent nodes the nominal semantic entities “.fibre” and / or “cable_cable_line”; the theme T2 comprising the sub-theme ST21 which has as parent node the nominal semantic entity “appareil_téléphone_tv_télévision”.

[0081] Thus, the implementation of steps 12, 13 and 14 allows the identification, from a set of verbatims, of typical actions representative of various operating contexts detected in the verbatims, and their automatic classification into sub-themes and themes. In other words, the present technique makes it possible to draw up, using an unsupervised approach, an overview of the different operating contexts identifiable in a set of verbatims, at different degrees of granularity, including at least a fine degree of granularity (i.e. categorization into typical actions) and more general degrees of granularity (categorization into sub-themes and themes).

[0082] In a particular embodiment, the analysis method according to the present technique also comprises, for example in parallel with steps 12, 13 and 14, a step 11 comprising the segmentation of the verbatims into text segments, and the determination of a polarity associated with each of said text segments. By polarity, we mean here an attribute representative of a feeling and / or an emotional state associated with the text segment considered. Such an attribute can for example qualify the expression within the text segment considered of a satisfaction or on the contrary of a dissatisfaction of a user, possibly to different degrees.

[0083] The segmentation of a verbatim is carried out automatically (i.e. without manual intervention), for example on the basis of punctuation characters present in the verbatim (e.g. commas, semicolons, periods, etc.). According to a particular characteristic, the segmentation also takes into account predefined text separators, including for example keywords classically used to introduce a notion of contradiction (e.g. the words "however", "on the other hand", "but", "nevertheless", "on the other hand", etc.).

[0084] At the end of step 11, each text segment is thus associated with a polarity or category of feeling among for example the following categories: a negative feeling (i.e. the expression of dissatisfaction), a neutral feeling (i.e. neither satisfied nor dissatisfied), a positive feeling (i.e. the expression of satisfaction). According to a particular characteristic, the determination of the polarities associated with the text segments is based on the use of a BERT language model (from the English “Bidirectional Encoder Representations from Transformers”) pre-trained to carry out sentiment analysis operations.

[0085] In a particular embodiment, in a step 15, the typical actions obtained at the end of steps 12, 13 and 14 are weighted according to the polarities associated with the text segments as identified in step 11. In other words, to the extent that the proposed technique makes it possible to associate a text segment with a typical action on the one hand and with a polarity on the other hand, it is possible to establish a link between polarities and typical actions, that is to say between a user feeling and a reason (or cause) of this feeling. The proposed technique makes it possible more particularly to identify trends at a fine level of granularity, by taking into account the number of text segments within which a particular feeling is associated with a particular reason.

[0086] According to a particular characteristic, the typical actions being furthermore organized into automatically determined sub-themes and themes, the present technique also makes it possible in one embodiment to establish a tree structure (for example in the form of a graph) of the typical actions by sub-themes and themes according to each type of polarity identified.

[0087] Such a tree structure can in particular serve as a basis for the construction of a graphical interface allowing a user to explore at different levels of granularity, feelings and reasons for feelings automatically extracted from a set of verbatims, thus greatly facilitating exploitation. It can also be used to automatically generate various summary reports in connection with business applications (in particular in the field of customer relationship management), such as satisfaction reports for example. According to another aspect, the present technique also relates to an electronic device for analyzing a set of textual data, called a set of verbatims, this device being capable of implementing the method previously described in any one of its embodiments.More particularly, such an electronic device comprises: means for determining recurring word combinations in said verbatims, and for identifying common patterns within said recurring word combinations; means for identifying synonymous relationships between words of the same syntactic function within said common patterns, and for grouping said synonymous words of the same syntactic function within the same data structures, called semantic entities; means for determining typical actions, as a function of conditional relationships identified between said semantic entities within said verbatims, a typical action being representative of similar contexts of operations identifiable in said verbatims.

[0088] Figure 3 represents, in a schematic and simplified manner, the structure of such an electronic device, in a particular embodiment. In a particular embodiment, this electronic device takes for example the form of a processing server connected via a communication network to at least one remote data source within which verbatims are stored, or even a processing server itself hosting one or more business applications capable of generating verbatims.

[0089] The electronic device according to the proposed technique comprises for example a memory 31 consisting of a buffer memory M, a processing unit 32, equipped for example with a microprocessor pP, and controlled by the computer program Pg 33, implementing the analysis method according to the invention.

[0090] Upon initialization, the code instructions of the computer program 33 are loaded into the buffer memory before being executed by the processor of the processing unit 32. The processing unit 32 receives as input E a set of textual data, or set of verbatims, each verbatim comprising an ordered sequence of words.

[0091] The microprocessor of the processing unit 32 then carries out the steps of the method previously described, according to the instructions of the computer program 33, to deliver at output S data comprising at least one set of typical actions representative of similar contexts of operations identifiable in said verbatims. In a particular embodiment, the output S takes for example the form of a structured set of data in which the typical actions are grouped by sub-themes, and the sub-themes by themes, according to different types of associated polarities (or feelings, e.g. negative, neutral, or positive) identified in the set of verbatims.

Claims

CLAIMS 1. Method for analyzing a set of textual data, called a set (EV) of verbatims, each verbatim comprising an ordered sequence of words, said method being implemented by an electronic device, said method being characterized in that it comprises: a step (11) of segmenting said verbatims into textual segments, and determining a polarity associated with each of said textual segments; a step (12) of determining recurring word combinations in said verbatims, and identifying common patterns within said recurring word combinations; a step (13) of identifying synonymous relationships between words of the same syntactic function within said common patterns, and grouping said synonymous words of the same syntactic function within the same data structures, called semantic entities;a step (14) of determining typical actions based on said semantic entities, a typical action being representative of similar contexts of operations identifiable in said verbatims; a step (15) of weighting said typical actions, a typical action being weighted based on the polarities associated with the textual segments of the verbatims with which said typical action is associated.; 2. Analysis method according to claim 1 characterized in that said grouping into semantic entities further comprises taking into account derivation relationships and / or spelling corrections between words of said common patterns.

3. Method according to claim 1, characterized in that said semantic entities comprise semantic entities of the verb, noun, or adjective type, and in that a typical action comprises a semantic entity of the verb type and at least one semantic entity of the noun or adjective type.

4. Analysis method according to claim 3, characterized in that it comprises the association of said typical actions with said verbatims, a typical action being associated with a verbatim when a combination of words present in said verbatim is associated with a common pattern comprising at least one word associated with the semantic entity of verb type of said typical action and at least one word associated with said at least one semantic entity of noun or adjective type of said typical action.

5. Analysis method according to claim 1 characterized in that said determination of typical actions comprises the identification of conditional relationships between said semantic entities, based on a probability of coexistence of words of said semantic entities within the same verbatims.

6. Analysis method according to claim 5 characterized in that said determination of typical actions comprises the construction of at least one graph as a function of said conditional relationships, and the use of said at least one graph to group said typical actions into sub-themes, and said sub-themes into themes.

7. Electronic device for analyzing a set of textual data, called a set of verbatims, each verbatim comprising an ordered sequence of words, said device being characterized in that it comprises: means for segmenting said verbatims into textual segments, and for determining a polarity associated with each of said textual segments; means for determining recurring word combinations in said verbatims, and for identifying common patterns within said recurring word combinations; means for identifying synonymous relationships between words of the same syntactic function within said common patterns, and for grouping said synonymous words of the same syntactic function within the same data structures, called semantic entities; means for determining typical actions, based on conditional relationships identified between said semantic entities within said verbatims, a typical action being representative of similar contexts of operations identifiable in said verbatims; means for weighting said typical actions, a typical action being weighted based on the polarities associated with the textual segments of the verbatims with which said typical action is associated.

8. Computer program product downloadable from a communication network and / or stored on a computer-readable medium and / or executable by a microprocessor, characterized in that it comprises program code instructions for executing a method of analyzing a set of textual data according to any one of claims 1 to 6, when executed by a computer.