METHOD AND SYSTEM FOR ASSIGNING SCORE
Patent Information
- Application Number
- FR2008001661
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2008-03-26
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2028-03-26
AI Technical Summary
Existing technologies struggle to effectively process and analyze large volumes of unstructured information due to heterogeneity issues, making it difficult to integrate and synthesize diverse data types for decision-making and operational processes.
A method and system that transforms unstructured data into structured data through a scoring process, using interval subdivision and modeling to assign scores to elements, enabling efficient analysis and storage in a database management system.
Enables effective analysis and management of large volumes of unstructured data by converting it into structured form, enhancing decision-making and operational processes by integrating heterogeneous data types.
Smart Images

Figure 00000015_0000 
Figure 00000016_0000 
Figure 00000017_0000
Abstract
Description
Method and system for assigning scores The invention relates to a method and a system for assigning a score to an element in a set of structured data derived from structured and / or unstructured data (texts, images, speech, etc.), the unstructured data being transformed into structured data via an automatic analysis process. Based on this score, an operational system can trigger an alert or initiate other control actions resulting from the analysis according to the invention.In the following description, the term "structured" refers to a set of data that can be represented as a relational table, represented by a rectangular matrix R of dimensions N x M, where N and M represent, respectively, the number of rows and the number of columns in matrix R. The invention is particularly applicable in the field of automatic parameter analysis of unstructured information when the volumes of information are large and cannot be analyzed manually. It can be implemented for web pages, emails, word processing documents, multimedia files, video files, text files, etc. From a technical point of view, the invention provides a solution to the problems of synthesizing heterogeneous information and managing unstructured information.25 Unstructured information now represents the majority of data collected and used in the professional world. It plays an important role in the management of business processes, but the fact that unstructured information is not immediately available and usable within the framework of these processes constitutes a major obstacle. 30 The coding of unstructured information and the model used in this coding (dimensions or parameters) allows for storage in a system of... Database management systems (DBMS) make unstructured information available and usable by business processes (decision-making / operational processes). In many sectors, companies generate, store, and manage large amounts of information in electronic form. Access to and understanding of this information can play a significant role in decision-making at all levels of the organization (marketing strategy, sales strategy, quality management, customer relationship management, etc.). This information is, in most cases, presented in an unstructured format that hinders easy content analysis. Given the sheer volume of this information, most organizations rely on automated text analysis techniques. Various methods are known in the prior art for resolving technical problems arising in the field of automatic text analysis. This automatic analysis focuses, for example, on sentiment and opinion analysis, risk analysis, and so on. One such technique is multimodal fusion. Companies today need methodologies that allow them to automatically synthesize information of different types: text and structured data, speech and structured data, text and speech, etc. For example, in the field of customer relationship management (CRM), companies need to connect information about customer needs and expectations (from phone calls, customer letters, customer messages or emails, surveys, forums, etc.).) and information derived from the analysis of behavioral and demographic data. This linking involves the integration and synthesis of heterogeneous, unstructured data such as data. Textual data, on the one hand, and structured data, on the other, are all relevant. Methods for processing heterogeneous information are also well-known. The issue of the heterogeneity of the data to be processed is linked not only to multimodality but also to the intrinsic heterogeneity of each type of data. For example: If we are interested in textual data from written sources, from which information such as feelings and opinions can be extracted, the user is faced with free-form texts—summaries of letters, emails, or customer transcripts of telephone calls, open-ended responses to opinion surveys—which consist of highly heterogeneous data in terms of source, nature, and quality, in the case of structured data, and in terms of source, nature or genre, quality, register, and idiom, in the case of unstructured data.From an automated analysis perspective, accounting for this structural heterogeneity is a methodological imperative that guarantees the efficiency and quality of the results obtained at the end of the analysis, whether it is conducted for decision-making and / or operational purposes. Discourse modeling methods also exist. Extracting sentiments and opinions from text streams or transcribed speech necessitates discourse modeling. US Patent 7,249,312 discloses a method and system for scoring unstructured data. The patent holder uses a maximum likelihood method to assign a score to portions of a document or stream, and then aggregates the scores to assign a final score to the document or data stream.One of the objectives of the present invention is to provide a method and a system enabling, in particular, the processing of large volumes of data. The invention is based in particular on the use of coding via a scoring process, that is, a process of assigning a score to an element or set of elements, without converting unstructured information into structured information, during normal operation of the process. It also uses modeling and extraction steps for unstructured information, in order to analyze the content of texts, extract relevant information for the intended applications, and represent it in a structured form. A method for assigning a score to a selected parameter within a database, characterized in that it comprises at least the following steps: let EG be the number of parameters retained to analyze the information base, and B: the maximum frequency of occurrence of one of the 15 parameters, - subdivide the interval [0, 1] into EG equal intervals Ik of width 1 / EG - define for 1 k -. EG of the intervals Ik = C(k - 1), k C defined from the L EG EG L as follows: to each interval Ik for 1 k (EG - 1) 20 associate an integer-valued interval Ik such that: Ik = [B(k-'),Bk[ a database to be analyzed being represented by the dimension vector EG such that: D, =(a,,a2,...,aEG) where 0 <=a,, (B-1) is the frequency of a parameter G,, in the database 25 of information, then > determine the decimal value V,.G of the overall score for this parameter G,, using the formula: EG VG =Ea xB(u-' if V,.G belongs to the integer-bounded interval Ike = [me,Me[ corresponding to the integer-valued interval Ik = [md,Md[ with decimal values, determine the overall score NG for the considered parameter of the database in using the relation: EG x (Me ∈ me) Me and me are the bounds of the integer interval and Md,md the bounds of the decimal interval. It includes, for example, a step of transforming unstructured data into structured data. It may include a step during which a business process executed on a client workstation retrieves information from a database storing the data-score pairs, and, based on the data-score information, generates an alert. The last interval IEG is, for example, associated with the integer interval IEG = [B('.G->>, (BEG -1)[. It uses, for example, a text document as its database.The invention relates to a system for implementing the process described above, characterized in that it comprises at least the following elements: means for detecting the data to be analyzed, means for analyzing said data by executing the steps of the process described above, and means for storing the data to which a score is assigned. The data is, for example, unstructured data, and the system includes a module for formatting said data. The system comprises, for example, several client workstations executing operational and / or decision-making processes, and in that said processes retrieve information from said storage means in order to... Check the data pairs, score, and depending on the score, trigger or not an alert. Other features and advantages of the device according to the invention will become clearer from the following description of an example embodiment given by way of illustration and in no way limiting, annexed to the figures which represent: - Figure 1, an example of implementation of the process according to the invention, - Figure 2, a system architecture implementing the process, - Figure 3, an example of CRM modeling and Figure 4 a detail of Figure 3. In order to better understand the principles implemented in the process 15 according to the invention, the following example is given in the case of a linguistic analysis for which the goal is in particular to: - code unstructured information into structured information, when the latter does not come in a format directly usable by the scoring process according to the invention, 20 - calculate a positioning indicator in the opinion direction which will be associated with Sun client, - synthesize the data ù structured data and unstructured data ù by correlating this indicator with the structured information associated with the unstructured data analyzed, for example, 25 demographic data and behavioral data (for example in the banking context, a complaint addressed to the bank will be via the name or identifier of the client associatable with the client data of the IS (demographic data and behavioral data).After processing by the method, the customer's 30th complaint will be represented by a structured data set: the data relating to opinion or sentiment will be. Scored. This score will be associated with demographic and behavioral data, enriching customer knowledge and understanding. Considering a customer segmentation application, again within the banking sector, analyzing such profiles will allow for the identification of customer segments sharing demographic, behavioral, and opinion data. Figure 1 illustrates a flowchart of the steps involved in implementing the method according to the invention. The input data 1, which may include structured and unstructured data, is first subjected to preprocessing 2, the function of which is, in particular, to differentiate the information in the structured data 3 from the unstructured data 4. The unstructured data 4 then undergoes processing 5, 7, and 8, which yields data in a structured format that allows the application of the scoring method 6 according to the invention detailed below. For this purpose, the processing method uses a textual data model 7 and lexical and grammatical resources 8 known to a person skilled in the art. The initially structured data, or the data structured after processing, is then merged 9 according to a method known to a person skilled in the art before being analyzed by executing the steps of the scoring method according to the invention.The data to which the process has assigned a score are then memorized and stored in a database 10. Figure 2 schematically illustrates an example of a system according to the invention, comprising a part 30 including various information sensors, such as telephone platforms 20, messages from messaging services 21, surveys 22, mail 23, or any other means 24 for obtaining and capturing information. The various elements or information are transmitted to the analysis server 11, comprising the various modules described in Figure 1. The processed data from the analysis server are Data is transmitted to a database 25. This data is then transmitted to decision-making processes and / or operational processes executed on client workstations 26. The business processes retrieve the information they need to execute from the database 25. Examples of business processes are given below. Decision-making processes can be implemented as business processes. This corresponds to profiling techniques for targeting customers based on their personal data and their opinions on specific products or services, or their expectations; customer segmentation techniques that take into account both identifying and behavioral data describing customers and the opinions they have expressed on specific products or services, etc., can be used. As another business process, operational processes can also be used. For example, within a CRM or quality process, all incoming emails associated with a given dissatisfaction score can be automatically and systematically routed to a service that will process them as a priority. Within a crisis management process, alerts can be triggered when the severity score of events or when the citizen dissatisfaction score reaches a given threshold. Implementation of the Scorinq Method: The scoring method according to the invention consists, in particular, in the example given for illustrative purposes only and not as a limitation, of assigning scores to sentiments and / or opinions expressed in a given document. A first step, therefore, is to determine which data or parameters should be assigned a rating or score.These parameters can be represented as words or phrases in a document, but could also, without departing from the scope of the invention, refer to parameters from measurement sensors in an industrial application. Indeed, it would be possible to consider a set of parameters such as temperature, pressure, etc., for which a user wishes to determine their significance and influence on the progress of a process. Once the parameters have been determined, the steps of the scoring process unfold as described below. By using the following notation: EG = number of parameters used to analyze a document, or an industrial phenomenon, B: the maximum frequency of occurrence of one of the parameters, ~o The process subdivides the interval [0, 1] into EG equal intervals Ik of width 1 / EG. We then decide that for 1 k EG, which leads to defining intervals ((kù1) kr I k L EG ' EG, to each interval Ik for 1 k (EG -1) is associated an interval with integer values Ik such that: Ik = [B(k ", Bk [, for the last interval IEG is associated the interval with integer values IEG = [B(EG->>, (BEG -1)[, representing a document to be analyzed by the vector of dimension EG such that: 20 =(al aZ,...aEG) where 0 5 aä <_ (B -1) is the frequency of a parameter G,, in the document, then the decimal value V,.G of the overall score for this parameter Gä in the document will be given by: EG VG = Eau x B(,'->> u=1 25 if V,.G belongs to the integer-bounded interval Ik = [me,Me[, corresponding to the integer-bounded interval Ik = [md,Md[. For decimal values, the overall score N;G for the considered document parameter will be given by the relation: + VG ùme EG x (Meùme). The parameter considered in the document is, for example, the severity parameter. Data Modeling: In cases where the data to be processed is not initially in a form compatible with the scoriaceous method according to the invention, the process performs a pre-processing step on the initial data in order to put it into a structured data format. This pre-processing step relies on linguistic processing, dependent on the data modeling. The modeling principle used makes it possible, in particular, to take into account the business, the nature of the data, and the analysis needs for a given application. It is similar to database models and gives the approach operational value.Figure 3 illustrates the various steps involved in modeling a discourse. Regardless of the domain being addressed and the nature of the data to be processed, the method relies on modeling the content of the analyzed unstructured information in the form of dimensions or parameters. This modeling approach allows for the definition of a Model component whose structure produces results that can be directly integrated into a database management system. Figure 3 illustrates an analytical model defined for the field of customer relationship management, or CRM (Customer Relationship Management), or customer understanding. This model allows, in particular, the analysis of customer / company interactions such as telephone calls, emails, complaint letters, and satisfaction surveys with open-ended questions. The parameters used to illustrate the process according to the invention are, for example, the following: - The dimension of facts (FACTS) 30, which allows for the identification of the statement of the problem encountered; - The dimension of feelings (FEELINGS) 31, which allows for the identification of the speaker's particular viewpoint on the axis of affective values; - The dimension of OPINIONS (OPINIONS) 32, which allows for the identification of the speaker's particular viewpoint on the axis of intellectual values. Depending on the level of analysis required and the application needs, this four-dimensional structure can be enriched with new dimensions, such as the dimension of needs (NEEDS) 33 or expectations 34 expressed by customers, the dimension of competitors (COMPETITORS), etc. Figure 4 highlights and details certain dimensions of Figure 3.For example, for the dimensions FEELINGS 31 and OPINIONS 32, it is possible to define three sub-dimensions which maintain a hierarchical relationship of dependence: 20 The first represents the dimension of polarities with three values: Negative, Neutral, Positive. Under the dimension of polarities M1 is the dimension of semantic classes M2, which allows feelings and opinions to be organized into classes according to a predefined typology. 25 Under the dimension of semantic classes is the dimension of degrees M3, which allows the semantic classes to be sub-categorized into three values, according to a criterion of degree or intensity. The example relating to an extract from a complaints database is given by no means as an exhaustive illustration of this approach.30 Unstructured input data comes from a manual transcription of telephone calls, by call center agents, for example: the customer is extremely disappointed by the new account statements... The model and the information extraction tools associated with it will allow the text extract to be annotated as follows: The customer is [extremely disappointed] feeling / negative / disappointment / degree 5 3 Some complex texts may require a degradation of the model which, for example, may be of the type (FEELINGS / Axis of polarities / Semantic Classes+ OPINIONS / Axis of polarities / Semantic Classes). 10 The system and method according to the invention described above can be implemented and applied in numerous fields, for example: for Intelligence Actors (of all types), CRM Actors, Quality Management, Process Control and Monitoring within a factory, for example, weighting of different measured parameters. They are also useful 15 for Risk Analysis in general: health risk, environmental risk, etc. ...crisis management, intelligence, research firms, consulting firms, and suppliers of products and services. 20 13
Claims
CLAIMS 1 - Method for assigning a score to a parameter selected from a database, the parameter coming from a measurement sensor and the method being characterized in that it comprises at least the following steps: let EG be the number of parameters retained to analyze the database, and B: the maximum frequency of appearance of one of the parameters, > subdivide the interval [0, 1] into EG equal intervals Ek of width 1 / EG > set for \ <k<EG intervals (£-1) k EG ' EG defined in the manner following: to each interval Ek for \< k <(EG -\) associate an interval with integer values Iek such that: Iek = [5ai),B*[ a database to be analyzed being represented by the dimension vector EG such that: where 0 <aa <(B-1) est la fréquence d’un paramètre Gu dans la base de données, alors > determine the decimal value VG of the overall rating for this Gu parameter using the formula: EG the U H—1 if VG belongs to the interval with integer bounds = corresponding to the interval with integer values Irk = [ma,Ma[ with decimal values, determine the overall score NG for the parameter considered in the database using the relation: v'-m- Me and me are the limits of the interval with integer values and Md, md the limits of the interval with decimal values, a step during which a business process executed on a client workstation (26) will search for information in a database storing the pairs data, score, and that based on the information given, score generates an alert. 2 - Method according to claim 1 characterized in that the data is in the form of unstructured data and in that it comprises a step of transforming unstructured data into structured data. 3 - Method according to claim 1 characterized in that for the last interval 7' is associated the interval with integer values I" = \BGG '\(BiG -1)[ 4- Method according to claim 1 characterized in that it uses a text document as an information base. 5 - System for implementing the method according to claim 1 characterized in that it comprises at least the following elements: means for detecting (20, 21, 22) the data to be analyzed, means for analyzing (11) said data by executing the steps of the method according to claim 1, means for storing (25) the data to which a score is assigned and one or more client stations (26) executing operational and / or decision-making processes and in that said processes will search for information in said storage means in order to control the pairs (data, score) and depending on the score trigger or not an alert. 6- System according to claim 5 characterized in that the data is unstructured data and in that it comprises a module (5, 7, 8) for formatting said data.