Information processing methods, electronic devices, and memory

JP7899134B2Active Publication Date: 2026-08-03NEC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NEC CORP
Filing Date
2023-07-20
Publication Date
2026-08-03

Smart Images

  • Figure 0007899134000001
    Figure 0007899134000001
  • Figure 0007899134000002
    Figure 0007899134000002
  • Figure 0007899134000003
    Figure 0007899134000003
Patent Text Reader

Abstract

To provide an information processing method, a device, an installation, and a memory that extract keywords from a text set related to a target object to determine a target element for the target object.SOLUTION: A method includes: extracting a plurality of keywords from a non-structured text set for a target object; grouping at least part of the plurality of keywords based on the semantics of the plurality of keywords; and determining a target element corresponding to the keywords in one group based on a result of the grouping. The target element represents one side of the target object. A new element effecting the target object is thus specified.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Exemplary embodiments of the present disclosure generally relate to the field of computers, and more particularly, to methods, devices, equipment, and computer-readable memories for information processing.

Background Art

[0002] Using unstructured text can provide comments on products, services, etc. For example, user comments are often displayed on product purchase pages or service display pages. As another example, questionnaires include open-ended questions for respondents to provide text comments. Such unstructured text contains rich information about the object to be described, and we want to interpret and utilize such information.

Summary of the Invention

Means for Solving the Problems

[0003] In a first aspect of the present disclosure, an information processing method is provided. The method includes extracting a plurality of keywords from an unstructured text set for a target object, grouping at least some of the plurality of keywords based on the meanings of the plurality of keywords, and determining a target element corresponding to a group of keywords based on the result of the grouping, where the target element represents an aspect of the target object.

[0004] A second aspect of the present disclosure provides an electronic device, the electronic device including at least one processing circuit, the at least one processing circuit configured to extract a plurality of keywords from an unstructured text set for a target object, to group at least some of the plurality of keywords based on their meaning, and to determine a target element corresponding to one group of keywords based on the grouping result, the target element representing one aspect of the target object.

[0005] A third aspect of the present disclosure provides an electronic device, the device comprising at least one processing unit and at least one memory, the memory of which is coupled to the processing unit and stores instructions for execution by the at least one processing unit. The instructions are a method for causing the device to perform a first aspect when executed by the at least one processing unit.

[0006] A fourth aspect of this disclosure provides a computer-readable memory which stores a computer program thereon, which is executable by a processor to carry out the method of the first aspect.

[0007] A fifth aspect of this disclosure provides an information processing method, which includes: obtaining a group of target elements of a target object, determined based on an unstructured text set relating to the target object, where each target element represents an aspect of the target object; and determining at least one key element for the target object based on the group of target elements and a group of structured elements of the target object, wherein at least one of the target elements in the group is different from the group of structured elements.

[0008] A sixth aspect of this disclosure provides an electronic device, the electronic device comprising at least one processing circuit, the at least one processing circuit comprising: taking a group of target elements relating to a target object, the group of target elements being determined based on an unstructured set of text relating to the target object, each target element representing an aspect of the target object; determining at least one key element relating to the target object based on the group of target elements and a group of structured elements relating to the target object, at least one of the target elements within the group of target elements being different from the group of structured elements.

[0009] A seventh aspect of the present disclosure provides an electronic device comprising at least one processing unit and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When the instructions are executed by the at least one processing unit, the device causes the device to perform the method of the fifth aspect.

[0010] An eighth aspect of this disclosure provides a computer-readable memory which stores a computer program executable by a processor to carry out the method of the fifth aspect.

[0011] A ninth aspect of this disclosure provides an information processing method, which includes: obtaining a group of target elements of a target object, where the group of target elements is determined based on an unstructured text set relating to the target object, where each target element represents an aspect of the target object, and presenting an information gathering sheet for collecting a description of the target object, where the at least one key element is determined from a group of structured elements and a group of target elements of the target object.

[0012] A tenth aspect of the present disclosure provides an electronic device, the electronic device comprising at least one processing circuit, the at least one processing circuit configured to acquire a group of target elements of a target object, the group of target elements being determined based on an unstructured text set relating to the target object, each target element representing a manifestation of the target object, and an information gathering sheet being presented for collecting a description of the target object based on at least one key element of the target object, the at least one key element being determined from a group of structured elements and a group of target elements of the target object.

[0013] An eleventh aspect of the present disclosure provides an electronic device, the device including at least one processing unit and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions cause the device to perform the method of the seventh aspect when executed by the at least one processing unit.

[0014] A twelfth aspect of this disclosure provides a computer-readable memory that stores a computer program executable by a processor to perform the method of the ninth aspect.

[0015] It should be understood that nothing described in this disclosure is intended to limit or restrict any significant features of the embodiments of this disclosure, nor is it used to limit the scope of this disclosure. Other features of this disclosure will be readily apparent from the following description. [Brief explanation of the drawing]

[0016] The above-mentioned features and other features, advantages, and aspects of the embodiments of this disclosure will become clearer upon reference to the attached drawings and the following detailed description. In the drawings, identical or similar reference numerals represent identical or similar elements: [Figure 1]This is a schematic diagram of an exemplary environment in which the embodiments of this disclosure are implemented. [Figure 2] This is a schematic diagram of an example of an information collection sheet according to some embodiments of the present disclosure. [Figure 3] This is a flowchart of the process for determining a target element according to some embodiments of the present disclosure. [Figure 4] This is a schematic diagram of keyword grouping according to some embodiments of the present disclosure. [Figure 5] This is a schematic diagram of information related to target elements according to some embodiments of the present disclosure. [Figure 6A] This figure shows an example of measuring a target element according to some embodiments of the present disclosure. [Figure 6B] This figure shows another example of measuring a target element according to some embodiments of the present disclosure. [Figure 7] This is a flowchart of the process for determining key elements according to some embodiments of the present disclosure. [Figure 8A] This is a schematic diagram illustrating the selection of key elements from target elements and structured elements, respectively, according to some embodiments of the present disclosure. [Figure 8B] This is a schematic diagram illustrating the centralized selection of key elements from both target elements and structured elements, according to some embodiments of the present disclosure. [Figure 9] This figure shows a flowchart of the process for presenting a list of mobile phone information according to some embodiments of the present disclosure; [Figure 10] This is a schematic diagram of an updated version of the information gathering sheet according to some embodiments of the present disclosure. [Figure 11] This is a schematic diagram of hints regarding target elements according to some embodiments of the present disclosure. [Figure 12A] This is a schematic diagram of a machine learning model for propensity scores according to some embodiments of the present disclosure. [Figure 12B] A schematic diagram of a machine learning model for conditional outcome expectations according to some embodiments of the present disclosure, and [Figure 13] A block diagram of a device implementing multiple embodiments of the present disclosure.

Embodiments for Implementing the Invention

[0017] Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the accompanying drawings, the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described in the text. Rather, it should be understood that these embodiments are provided to more fully and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are used only for illustrative operations and not for limiting the scope of protection of the present disclosure.

[0018] In the description of the embodiments of the present disclosure, the terms "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based" should be understood as "at least partially based". The term "one embodiment" or "this embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions can also be included below.

[0019] As used in this text, the term “circuit” means a hardware circuit and / or a combination of a hardware circuit and software. For example, a circuit may be a combination of analog and / or digital hardware circuitry and software / firmware. Another example is a circuit, which is any part of a hardware processor having software that includes (multiple) digital signal processors and (multiple) memories that work together to enable the device to operate to perform various functions. Yet another example is a hardware circuit and / or processor, such as a microprocessor or a part thereof, which requires software / firmware for operation, but may not have software if it is not required for operation. As used in this text, the term “circuit” includes hardware circuitry or processors only, or parts of hardware circuitry or processors, as well as implementations of the software and / or firmware associated with them (or them).

[0020] As used in this text, the term "model" refers to a system that can learn associations between corresponding inputs and outputs from training data and, after training is complete, generates a corresponding output for a given input. Model generation is based on machine learning techniques. Deep learning is a machine learning algorithm that uses multiple layers of processing units to process inputs and provide corresponding outputs. In this text, "model" may also be referred to as a "machine learning model," "machine learning network," or "network," and these terms are used interchangeably. A model can further include different types of processing units or networks.

[0021] As mentioned above, unstructured text about a subject contains a wealth of information about that subject. It is expected that this information will be interpreted and utilized. In conventional scenarios, the subject is described using manually specified sentences or sentences extracted from other texts. This conventional method fails to extract elements related to the subject, nor can the degree of concern of those elements be quantified. Therefore, conventional schemes have limitations in interpreting unstructured text and do not provide information for further utilization.

[0022] Embodiments of this disclosure propose a method for information processing. In one embodiment of this disclosure, keywords are extracted from a text set relating to a target object, and a group of target elements for the target object is determined based on the grouping of the extracted keywords. Each target element represents a manifestation of the target object. By extracting target elements from unstructured text, new elements that affect the target object are discovered.

[0023] In another aspect of this disclosure, at least one key element of a target is determined from a group of target elements and a group of structured elements for the target. At least one target element is distinct from the structured elements. By considering both existing structured elements and newly extracted target elements, the key elements are determined more accurately. This helps in recognizing key aspects of the target and facilitates the optimization of the target.

[0024] In yet another aspect of this disclosure, an information gathering sheet is presented for collecting descriptions of a target subject based on at least one determined key element. By optimizing the design of the information gathering sheet in this way, evaluation information about the target subject is collected more efficiently.

[0025] Sample environment Figure 1 is a schematic diagram of an exemplary environment 100 in which embodiments of the present disclosure are carried out. In environment 100, a first computing device 110 receives a text set 105 relating to a target object, or the first computing device 110 extracts a text set 105 from original data. The text set 105 includes a plurality of texts 101-1, 101-2, ..., which are collectively or individually referred to as text 101. The target object includes tangible objects, intangible objects, and combinations thereof. For example, the target object may be products such as household goods and food. Another example is that the target object may be a service such as a cloud computing service or a cloud storage service. Another example is that the target object may be an entity that provides services or goods, such as an airline, a restaurant, or a hotel.

[0026] Text 101 may be a description of the target subject by the target user. Text 101 may include emotional sentences such as "Apples are delicious" or "Apples are not delicious." Text 101 may also include non-emotional sentences such as "I ate an apple." Text 101 may also include evaluations, comments, reviews, assessments, advice, and impressions of the target subject. Text 101 may include information about factors that influence the target subject. Individual texts 101 within text set 105 may be provided by different users or by the same user at different times.

[0027] In some embodiments, text 101 may be a user's rating on the target display page. The display page may originate from, for example, a shopping app, a service provider app, or a review app.

[0028] In some embodiments, as shown in Figure 1, text 101 may be generated from an information gathering sheet 150 for the target subject. As used in this text, the "information gathering sheet" is for collecting descriptions (e.g., evaluations, impressions, etc.) about the target subject, and may be, for example, an electronic questionnaire or comments. The information gathering sheet 150 includes open-ended questions about the target subject. Text 101 may be the user's answers to the open-ended questions.

[0029] Figure 2 shows an example of an information gathering sheet 150. In this example, the information gathering sheet 150 for a particular flight includes an open question 230. Users can provide their evaluation of the flight, etc., through a text box. The response set 250 of the information gathering sheet 150 is shown in tabular format. Each row in the response set 250 represents a response record from the same user. In each response record, column 258 is the answer to the open question 230. Text 101 may also be the text in column 258.

[0030] Refer to Figure 1 again. The first computing device 110 determines the target elements 102-1, 102-2, ... which are also collectively referred to as a group of target elements 102, or simply as target elements 102, based on the text set 105. Since such target elements 102 are determined from unstructured text, they are also called "extracted elements" or "unstructured elements".

[0031] A group of target elements 102 is provided to a second computing device 120. The second computing device 120 also receives or determines target structured elements 103-1, 103-2, ... which are also collectively referred to as a group of structured elements 103, or separately referred to as structured elements 103. As used in this text, the term “structured element” means an element whose criterion has predetermined options (e.g., predetermined numerical value, category, star rating, etc.). For structured elements, the user evaluates or describes the target from the perspective of that structured element by selecting one option from a predetermined set of choices. Structured elements are quantitative and highly organized. The description of structured elements (e.g., evaluation, assessment) is not open and must conform to an architecture with predetermined options.

[0032] Structured elements include numerical elements or categorical elements. The given options for numerical elements include predetermined numbers, stars, etc. The given options for categorical elements include predetermined classes, such as cabin classes. In this text, target elements and structured elements are collectively referred to as "elements," or individually.

[0033] In some embodiments, a group of structured elements 103 may be obtained from an information gathering sheet 150, as shown in Figure 1. The information gathering sheet 150 includes closed questions about the structured elements 103. A "closed question" is a question in which the answer is selected from a predetermined set of choices. In the example in Figure 2, the information gathering sheet 150 includes closed question 210-1 about the structured element "seat comfort", closed question 210-2 about the structured element "cabin service", closed question 210-3 about the structured element "food and beverage", closed question 210-4 about the structured element "entertainment", closed question 210-5 about the structured element "ground service", and closed question 210-6 about the structured element "value for money". Closed questions 210-1 to 210-6 are collectively referred to as closed questions 210, or are referred to individually. Each closed question 210 has five scores for the user to select. In response set 250, columns 252 to 257 are the user's responses to closed questions 210-1 to 210-6, respectively.

[0034] Refer to Figure 1 again. The second computing device 120 determines at least one key element 104-1, 104-2, ... of a group of targets from a group of target elements 102 and a group of structured elements 103, the key element 104 is also called key element 104 collectively or individually. As used in the text, the term “key element” means an element that affects the target. The impact on the target includes the impact on the performance, service, functionality, overall evaluation, or satisfaction of the target. In particular, a key element may be an element that has a high degree of impact on the target out of many elements. The degree of impact reflects the importance of the element to the target.

[0035] In the example in Figure 2, the information gathering sheet 150 includes a closed question 220 regarding the overall assessment of the target group. In the response set 250, column 251 is the user's response to the closed question 220.

[0036] Refer to Figure 1 again. The determined key element 104 is provided to the third computing device 130. The third computing device 130 presents an information gathering sheet 160 for the target based on the key element 104. In some embodiments, the information gathering sheet 160 may be generated based on the key element 104. In some embodiments, the information gathering sheet 160 may be an updated version of the information gathering sheet 150.

[0037] In environment 100, the first computing device 110, the second computing device 120, and the third computing device 130 may be any type of computing device, including terminal devices or service end devices. Terminal devices are any type of mobile, fixed, or portable terminal, including any combination of the above, such as mobile handsets, desktop computers, laptop computers, netbooks, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / video cameras, positioning devices, television receivers, radio receivers, e-book devices, game devices, or components and peripherals of these devices. Service end devices include, for example, computing systems / servers such as mainframes, edge computing nodes, and computing devices in a cloud environment.

[0038] It should be understood that the structure and function of environment 100 are described for illustrative purposes only and do not imply any limitation to the scope of this disclosure. Although Figure 1 shows the first computing device 110, the second computing device 120, and the third computing device 130 separately, in some embodiments, both or all of the first computing device 110, the second computing device 120, and the third computing device 130 may be the same device or belong to the same computing system.

[0039] Furthermore, the information gathering sheet shown in Figure 2 is illustrative and not intended to limit the scope of this disclosure. The open-ended and closed-ended questions, and their number, shown in Figure 2 are illustrative. In embodiments of this disclosure, the information gathering sheet may have any appropriate number of open-ended and closed-ended questions. Furthermore, while English is given as an example, embodiments of this disclosure may be used to process text and information gathering sheets in any language.

[0040] Extraction of target elements Figure 3 is a flowchart of a process 300 for determining a target element according to several embodiments of the present disclosure. Process 300 is performed in a first computing device 110. For ease of explanation, process 300 will be described with reference to Figure 1.

[0041] In block 310, the first computing device 110 extracts multiple keywords from an unstructured text set 105 for the target. The extracted keywords have any appropriate number of word segments. The keywords include single-word keywords such as "flight," "seat," and "service," and two-word keywords such as "cabin crew" and "flight attendant." Any appropriate keyword extraction algorithm can be used, but is not limited to, TF-IDF, KP-Miner, SBKE, RAKE, TextRank, YAKE, KeyBERT, etc.

[0042] In some embodiments, the text 101 in the text set 105 can be preprocessed before applying the keyword extraction algorithm, for example, by removing named entities and stop words. Named entities are, for example, names of people, organizations, places, etc., that do not describe any aspect of the target. In the case of English text, stop words are, for example, "a," "an," "the," and "and." In the case of Chinese text, stop words are, for example, "one," "one," "and," and "however." Alternatively, in some embodiments, the text 101 may be preprocessed by the keyword extraction algorithm.

[0043] In some embodiments, a keyword extraction algorithm is used to extract nouns as keywords from the text set 105. This avoids extracting words that cannot describe other attributes of the target aspect, thereby effectively reducing the difficulty of subsequent processing.

[0044] In some embodiments, the first computing device 110 extracts keywords based on the number of occurrences (i.e., word frequency) of each word in the text set 105. Specifically, the first computing device 110 extracts candidate words from the text 101 of the text set 105. If the number of occurrences of a candidate word in the text set 105 is greater than a threshold number, the candidate word is determined to be one of the keywords. If the number of occurrences of a candidate word in the text set 105 is less than a threshold number, the candidate word is deleted.

[0045] For example, a keyword extraction algorithm is used to extract candidate words from column 258 of each answer record. For each extracted candidate word, the number of occurrences of that word in the entire text set 105 is calculated. Then, candidate words whose occurrence count is greater than a threshold count are identified as keywords, and candidate words whose occurrence count is less than the threshold count are removed. In such an embodiment, filtering the pre-extracted candidate words prevents unimportant words from interfering with the determination of the target element.

[0046] Alternatively, in some embodiments, the first computing device 110 extracts keywords based on the meaning of the text 101 in the text set 105. For example, semantic analysis determines sentences that have emotion, and nouns related to emotion in such sentences are used as keywords.

[0047] In block 320, the first computing device 110 groups at least some of the multiple keywords based on the meaning of the multiple keywords. In some embodiments, all keywords are grouped. In some embodiments, the keywords are filtered based on the results of preliminary grouping, and the filtered keywords may be grouped.

[0048] The first computing device 110 uses a cluster to group the extracted keywords. Therefore, it generates a word vector representing the meaning for each keyword. The word vector is generated using any suitable method, such as word2vector or GloVe. Embodiments of this disclosure are not limited in this respect.

[0049] Multiple keywords can be clustered based on word vectors to determine multiple clusters, each cluster containing at least one keyword. The clustering algorithm divides these keywords into independent, non-overlapping clusters based on their semantic similarity. Any suitable clustering algorithm can be employed, such as K-means, Density-Based Spatial Clustering of Applications with Noise (DBSCAN), or Gaussian mixture models.

[0050] In some embodiments, keywords may be filtered based on the quality of each cluster. Cluster quality represents the semantic degree to which the keywords within that cluster are grouped together. For example, the sum of the squared distances of the keywords within a cluster can be used as the cluster quality. Alternatively or additionally, a silhouette coefficient can also be used as the cluster quality.

[0051] The quality of each cluster obtained through clustering is determined. In some embodiments, keywords in clusters with a quality lower than the threshold quality are removed to determine the remaining keywords. The remaining keywords are grouped based on their meaning. For example, the remaining keywords are clustered. Keywords within the same resulting cluster are considered as a single group of keywords. Alternatively, in some embodiments, clusters with a quality lower than the threshold quality are removed, and other clusters with a quality higher than the threshold quality are retained. In reserved clusters, keywords within the same cluster are treated as a single group of keywords. In such embodiments, it is not necessary to regroup the remaining keywords.

[0052] Figure 4 shows an example of keyword grouping. The grouped results are obtained by processing the text in column 258 within the response set 250. In Figure 4, keyword groups 410, 420, 430, 440, 450, 460, and 470 are determined by clustering. Each keyword group contains one or more keywords. Refer to Figure 3 again. In block 330, the first computing device 110 determines a target element 102 corresponding to a group of keywords based on the grouping results. The target element 102 represents one aspect of the target object. Since keywords in the same group have similar meanings, they represent the same aspect of the target object. Thus, a group of keywords corresponds to one target element 102.

[0053] The name or identification of a target element 102 corresponding to a group of keywords is determined based on the group of keywords. For example, one of the group of keywords is used to represent the corresponding target element. Another example is determining the center of a cluster of keywords from the group and representing the corresponding target element with the keyword having the semantic characteristics closest to that center. A further example is that the target element is represented by the aspect of the target (e.g., service or performance) described by the group of keywords.

[0054] In the example in Figure 4, the target element corresponding to keyword group 410 is "tv service". The target element corresponding to keyword group 420 is "boarding procedure". The target element corresponding to keyword group 430 is "luggage service". The target element corresponding to keyword group 440 is "movie service". The target element corresponding to keyword group 450 is "price". The target element corresponding to keyword group 460 is "time". The target element corresponding to keyword group 470 is "legroom".

[0055] In some embodiments, one or more groups of keywords that are identical or similar to a structured element are removed. In this case, the first computing device 110 determines the target element corresponding to the group of keywords that were not removed. For example, for each group of keywords, the first computing device 110 determines whether the group of keywords is semantically similar to the target structured element. If the group of keywords is not semantically similar to any of the structured elements, the target element is determined based on the group of keywords. If the group of keywords is semantically similar to a structured element, the group of keywords is deleted.

[0056] As an example, processing the text in column 258 yields the set of keywords "food," "meal," "drink," and "snack." This keyword group is semantically similar to the structured element "food and beverage" in Figure 2. Therefore, one group of keywords is removed without determining its corresponding target element.

[0057] The above process 300 extracts the target elements from open text comments or comments. Analyzing the information contained in unstructured text in this way reveals new elements that influence the target.

[0058] Process 300 further includes additional blocks. In some embodiments, the first computing device 110 determines, based on the text set 105, at least one target sentence corresponding to the target element 102. The target sentence reflects a perspective on the target element. For example, the target sentence may be a clear sentence related to the target element. The target sentence can be used to interpret the target element.

[0059] Each target sentence must include the target element (e.g., describe or discuss) with clear emotion. The target sentence may reflect a positive perspective on the target element. Alternatively or additionally, the target sentence may reflect a negative view of the target element. Furthermore, each target sentence must be valid and understandable. In some embodiments, the target sentence includes only the target element and not other elements of the target subject. In such embodiments, the target sentence interprets a single element clearly to avoid confusion.

[0060] The first computing device 110 determines the target sentence in any appropriate way. For example, it generates one or more sentences related to the target element. It determines whether there are sentences in the text set 105 that are semantically similar to the generated sentence, and how many such sentences there are. If the number of matching sentences exceeds a threshold number, the generated sentence can be used as the target sentence.

[0061] In some embodiments, the first computing device 110 determines a target sentence using keywords corresponding to the target element. Specifically, the first computing device 110 extracts at least one candidate sentence from the text set 105. Each extracted candidate sentence contains at least one keyword from a group of keywords corresponding to the target element. For example, a candidate sentence extracted for the target element "leg room" contains at least one keyword from keyword group 470.

[0062] The first computing device 110 further determines at least one target sentence related to the target element based on the extracted candidate sentence. For example, the extracted candidate sentence may be used as the target sentence as is. Other examples include merging candidate sentences with the same emotion into a single target sentence, or generating a single target sentence based on candidate sentences with the same emotion.

[0063] The target sentences include those that reflect a positive view of the target element, those that reflect a negative view of the target element, or both, depending on the views about the target element reflected by the text in text set 105. That is, the target sentences include sentences with positive feelings, sentences with negative feelings, or both.

[0064] Table 500 in Figure 5 shows information about the target element "leg room". Target sentence 501 expresses a positive feeling towards the target element "leg room", while target sentence 502 expresses a negative feeling towards the target element "leg room". Presenting target elements individually can result in limited or unclear information. Using target sentences allows for interpretation of target elements. Presenting target sentences together with target elements allows stakeholders to understand the target elements more intuitively.

[0065] In some embodiments, the first computing device 110 can also determine the number of sentences in the text set 105 that have a similar meaning to the target sentence. This number can be used as the frequency of the target sentence. For example, Figure 5 shows that the frequency of target sentence 501 is 500 and the frequency of target sentence 502 is 800. This means that there are more negative views than positive views regarding the target element, "leg room." By determining and presenting the frequency of the target sentence, the relative importance of the target element can be intuitively grasped.

[0066] Quantification of target elements In some embodiments, the target element is further quantified. As used in this text, quantifying an element means determining a measurement of that element that represents the degree of attention, importance, or influence of that element. The measurement for the target element 102 is determined based on a group of keywords corresponding to the target element 102 and text 101 within a text set 105.

[0067] Such a measurement is expressed as the number of times a group of keywords corresponding to the target element 102 appears in the text 101. The measurement of the target element 102 is determined for each text 101. In this case, the number of occurrences of the keywords in each text 101 is determined as the measurement. In embodiments where the text 101 is generated from the information collection sheet 150, the number of occurrences of the keywords may be determined for each response record in the information collection sheet 150. In some embodiments, the sentiment of the text can also be analyzed, and the number of occurrences of keywords can be determined based on the sentiment of the text. Such embodiments are described below with reference to Figure 7.

[0068] For example, to determine the measurement of the target element "leg room" corresponding to keyword group 470, the number of occurrences of the keywords "leg," "leg room," and "leg space" in each text 101 is determined. Figure 6A shows an example of the measurement of the target element "leg room." Each number in column 610 represents the number of occurrences of the keywords "leg," "leg room," and "leg space" in the corresponding text in column 258.

[0069] Alternatively, the measurement for the target element 102 is expressed as the sentiment level of a sentence in text 101 that contains the keyword corresponding to the target element 102. The sentiment level is divided into, for example, five levels, each represented by a number from 1 to 5. The measurement for the target element 102 is determined for each text 101. In this case, the sentiment level of the sentence containing the keyword in each text 101 is determined. In embodiments where text 101 is generated from the information collection sheet 150, the sentiment level of the sentence is determined for each response record.

[0070] For example, taking the target element "leg room," to determine its measurement, the sentiment level of sentences containing at least one of the keywords "leg," "leg room," and "leg space" within each text 101 is determined. Figure 6B shows an example of measuring the target element "leg room." Each number in column 620 represents the sentiment level of sentences containing at least one of the keywords "leg," "leg room," and "leg space" within the corresponding text in column 258. A number "0" indicates that the corresponding text does not contain the keywords "leg," "leg room," or "leg space," or that sentences containing the keywords "leg," "leg room," or "leg space" are neutral sentences that are not emotional. Furthermore, the preliminary quantified values ​​are converted to match the measurement of structured elements. For example, a number "0" is converted to a score of 3, representing a neutral sentiment. The above is merely one example of target element quantification, and it should be understood that in embodiments of this disclosure, target elements may be quantified in any appropriate method and number of times.

[0071] The measured values ​​of the target elements shown in Figures 6A and 6B are illustrative and not intended to limit the scope of this disclosure. The quantification of the target elements described above is achieved by either or both of the first computing device 110 or the second computing device 120. The target elements can also be quantified using the methods described below.

[0072] Identifying Key Elements Figure 7 is a flowchart of a process 700 for determining key elements according to some embodiments of the present disclosure. Process 700 is performed on a second computing device 120. For ease of explanation, process 700 will be described with reference to Figures 1 and 7.

[0073] In block 710, the second computing device 120 acquires a target element 102 from a group of target objects. The target elements 102 are determined based on an unstructured text set 105 relating to the target objects, and each target element 102 represents one aspect of the target object.

[0074] In some embodiments, the second computing device 120 receives instructions for the target element 102 from the first computing device 110, as shown in Figure 1. Alternatively, in some embodiments, the second computing device 120 may determine the target element 102 based on the text set 105, as described above with reference to Figure 3.

[0075] The second computing device 120 also receives or determines a group of structured elements 103 to be targeted. The structured elements 103 originate from the information gathering sheet 150, as illustrated with reference to Figure 1, for example. At least one of the target elements 102 in the group is different from the group of structured elements 103. In some embodiments, all target elements 102 are different from the structured elements 103.

[0076] In block 720, the second computing device 120 determines at least one key element 104 for the target object based on a group of target elements 102 and a group of structured elements 103 of the target object. In some embodiments, the second computing device 120 determines the number of occurrences of the keyword corresponding to each element in the text set 105. The target elements 102 and structured elements 103 are ranked according to their frequency of occurrence, and a certain number of the top ones are determined to be key elements.

[0077] In some embodiments, to determine key elements, the second computing device 120 quantifies target elements and structured elements. Specifically, the second computing device 120 determines a first measurement for each of a group of target elements 102 by analyzing the sentiment of the texts 101 in the text set 105. The first measurement represents the attention given to the corresponding target element. The second computing device 120 can also determine a second measurement for each of a group of structured elements 103, representing the attention given to the corresponding structured element. To identify key elements, measurements of various types of elements should be performed consistently. Therefore, the first and second measurements are aligned with respect to the metrics scale.

[0078] In some embodiments, the first measurement may be expressed as the number of keyword occurrences. From the text 101 of the text set 105, for each target element in a group of target elements, a sentence containing the keyword corresponding to that target element and having sentiment is determined. The first measurement of a target element is determined based on the number of occurrences of the keyword corresponding to the target element within the sentence.

[0079] As an example, for the target element "leg room," sentences containing the keywords "leg," "leg room," and "leg space," and possessing emotion, are determined for each text (101). The number of times the keywords "leg," "leg room," and "leg space" appear in these sentences is determined as the first measure. For example, column 610 in Figure 6A shows the first measure for the target element "leg room."

[0080] In such embodiments, it is not appropriate to directly use the scores (ratings) in columns 252-258 as the second measurement in order to match the second measurement of the structured element with the first measurement. Therefore, it is necessary to requantify the structured element. Specifically, for each structured element in a set of structured elements, the second computing device 120 determines from text 101 sentences that contain the keyword corresponding to that structured element and have sentiment. The second measurement of the structured element is determined based on the number of occurrences of the keyword corresponding to the structured element within the sentence.

[0081] As an example, for the structured element "food and beverage," sentences containing the keywords "food," "meal," "drink," and "snack" and possessing emotion are identified in each text 101. The frequency of occurrence of the keywords "food," "meal," "drink," and "snack" in these sentences is determined as the second measure. For example, column 630 in Figure 6A shows the second measure for the structured element "food and beverage."

[0082] Alternatively, in some embodiments, the first measurement may be expressed as the sentiment level of a sentence containing a keyword. For each target element in a group of target elements, a sentence containing the keyword corresponding to that target element and having sentiment is determined from the text 101 in the text set 105. Based on the sentiment level of the sentence, the first measurement of the target element is determined. The sentiment level of the sentence can be determined in any suitable way, and embodiments of the present disclosure are not limited in this respect.

[0083] As an example, for the target element "leg room," sentences containing the keywords "leg," "leg room," and "leg space," and possessing emotion, are determined for each text (101). The emotion level of the sentence is determined as the first measurement. For example, column 620 in Figure 6B shows the first measurement for the target element "leg room."

[0084] In this embodiment, the second measurement of a structured element may be determined for each structured element in a group of structured elements 103 based on the answers to closed questions about that structured element. For example, the user's score for a structured element may be used as the second measurement. In Figure 6B, columns 252 to 258 are used as the second measurement for each structured element. As can be seen from Figures 6A and 6B, the first and second measurements are determined for each response record.

[0085] The determination of the first and second measurements has been explained above. The second computing device 120 further determines the degree of influence each element has on the target object based on the first measurement of each of the group of target elements 102 and the second measurement of each of the group of structured elements 103. The degree of influence is determined based on any appropriate algorithm. Such algorithms include, but are not limited to, linear regression, logical regression, and Sharley values.

[0086] Element strength is determined for each target element 102 and structured element 103 as an indicator of influence. Element strength represents the importance of the corresponding element to the results related to the target object. Results related to the target object include, for example, the performance of the target object, the overall evaluation of the target object, and satisfaction with the target object. In the example in Figure 2, the results related to the target object are the responses to the closed question 220, i.e., the scores listed in column 251.

[0087] Then, key elements are selected from the target elements 102 and structured elements 103 according to their degree of influence. For example, multiple elements with high influence can be selected. In this text, the process of ranking elements according to their degree of influence on the target (for example, the strength of the elements) is also called "key factor ranking (KFR)".

[0088] In some embodiments, key elements may be selected from a group of target elements 102 and a group of structuring elements 103, respectively. Specifically, a first number of target elements are selected as key elements from the group of target elements 102, depending on the degree to which the group of target elements 102 has an influence on the target object. A second number of structuring elements are selected as key elements from the group of structuring elements 103, depending on the degree to which each of the group of structuring elements 103 has an influence on the target object.

[0089] The values ​​of the first and second numbers may be predetermined. Alternatively, the selected element may be an element whose influence is greater than a threshold (for example, an element whose intensity is greater than the threshold intensity). In this case, the values ​​of the first and second numbers are not predetermined. The implementation of this disclosure is not limited in this respect.

[0090] Figure 8A shows the results of ranking key elements for both target elements and structured elements. The strength of the elements in the horizontal coordinates of Figure 8A indicates the degree to which the corresponding element influences the target. As shown in the figure, "price," "movie service," and "tv service" are selected as key elements from the target elements according to their strength. From the structured elements, the structured elements "value for money," "ground service," "cabin service," "seat comfort," and "food and beverage" are selected as key elements according to their strength.

[0091] In some embodiments, the key element may be selected from the set of the unions of a group of target elements 102 and a group of structured elements 103. Specifically, a third number of elements are selected as the key element from the union of a group of target elements 102 and a group of structured elements 103 with respect to the target object, depending on the degree of influence of the group of target elements 102 and the degree of influence of the group of structured elements 103 on the target object.

[0092] The value of the third number may be predetermined. Alternatively, the selected element may be an element whose influence is greater than a threshold (for example, an element whose intensity is greater than the threshold intensity). In this case, the value of the third number is not predetermined. The implementation of this disclosure is not limited in this respect.

[0093] Figure 8B shows the ranking results of key elements combined with target elements and structured elements. The element strength in the horizontal coordinates of Figure 8B represents the degree of influence that the corresponding element has on the target. As shown in the figure, the elements "value for money," "ground service," "cabin service," "seat comfort," "food and beverage," "price," "movie service," and "tv service" are selected as key elements according to their strength.

[0094] Presentation of information gathering sheet Figure 9 shows a flowchart of process 900 for presenting an information gathering sheet according to several embodiments of the present disclosure. Process 900 is performed on a third computing device 130. For ease of explanation, process 900 will be described with reference to Figures 1 and 9.

[0095] In block 910, the third computing device 130 acquires a target element 102 from a group of target objects. The target element 102 is determined based on an unstructured text set 105 relating to the target objects, and each target element 102 represents one aspect of the target object.

[0096] In some embodiments, as shown in Figure 1, the third computing device 130 receives instructions for the target element 102 from the first computing device 110. Alternatively, in some embodiments, the third computing device 130 may determine the target element 102 based on the text set 105, as described above with reference to Figure 3.

[0097] In block 920, the third computing device 130 presents an information gathering sheet for collecting a description of the target based on at least one key element 104 of the target. The at least one key element 104 is determined from a group of structured elements 103 and a group of target elements 102 of the target. In some embodiments, the third computing device 130 receives instructions from the second computing device 120 regarding at least one key element 104, as shown in Figure 1. Alternatively, in some embodiments, the third computing device 130 determines the key element from a group of structured elements 103 and a group of target elements 102, as described above with reference to Figure 7.

[0098] In some embodiments, the text 101 in the text set 105 is generated from answers to open questions in an information gathering sheet, which includes corresponding closed questions relating to a group of structured elements 103. In such embodiments, the third computing device 130 presents an updated version of the information gathering sheet based on at least one key element. This updated version includes updated closed questions. In this way, the new information gathering sheet more directly collects the user's assessment of areas of interest.

[0099] As an example, based on the information gathering sheet 150 and the corresponding answer set 250 shown in Figure 2, the key elements "value for money," "ground service," "cabin service," "seat comfort," "food and beverage," "price," "movie service," and "tv service" are identified, as shown in Figure 8B. The third computing device 130 presents an updated version of the information gathering sheet 150 shown in Figure 2, which is the information gathering sheet 160 shown in Figure 10. Comparing Figure 2 and Figure 10, closed questions 210-1 to 210-6 have been updated to closed questions 210-1, 210-2, 210-3, 210-5, 210-6, and 1010.

[0100] In some embodiments, if at least one key element 104 includes a target element, the third computing device 130 adds a closed question about the target element to the updated version of the information gathering sheet. Furthermore, the third computing device 130 presents the updated version of the information gathering sheet that includes the closed question.

[0101] Continuing the above example, the key element includes the target element "price". Therefore, the presented information gathering sheet 160 includes a closed question 1010 regarding the target element "price". In this way, aspects that users are likely to be interested in are added to the information gathering sheet as structured elements. This makes it possible to collect users' evaluations of the target product more comprehensively and easily.

[0102] In some embodiments, if at least one key element 104 does not contain a structured element, the third computing device 130 removes the closed question regarding the structured element from the information gathering sheet. Furthermore, the third computing device 130 presents an updated version from which the closed question regarding the structured element has been removed.

[0103] Continuing with the above example, the key elements do not include the structured element "entertainment" from Information Gathering Sheet 150. This eliminates the closed question 210-4 related to the structured element "entertainment." The presented Information Gathering Sheet 160 does not include the closed question 210-4 compared to Information Gathering Sheet 150. In this way, aspects that may not be of much interest to the user are removed from the Information Gathering Sheet. This avoids unimportant issues interfering with the user.

[0104] Alternatively, in some embodiments, the third computing device 130 may provide hints regarding target elements included in key elements while the information gathering sheet is being presented. These hints facilitate the user providing descriptions regarding target elements such as experience, evaluation, and satisfaction. Specifically, while the information gathering sheet 160 is being presented, the third computing device 130 detects answers to open-ended questions 230 within the information gathering sheet 160. If it is detected that answers have been provided, the third computing device 130 provides such hints.

[0105] Figure 11 shows an example of a hint regarding a target element. Continuing the example above, the main element includes the target element "movie service". As shown in the figure, the user enters the text "The food is OK, and" into the text box 1120 of the open question 230. In response to detecting that text has been entered, the third computing device 130 presents a hint 1110 "How about the movie" regarding the target element "movie service". Hint 1110 prompts the user to provide an experience or evaluation of the target element "movie service".

[0106] In some embodiments, the third computing device 130 interactively determines and presents an information gathering sheet 160. Specifically, the third computing device 130 presents at least one key element 104. While at least one key element is presented, the third computing device 130 detects the selection of at least one key element. If the selection of at least one key element is detected, the third computing device 130 adds a closed question regarding the selected key element to the information gathering sheet 160 and then presents the information gathering sheet 160 including the closed question.

[0107] In this embodiment, the third computing device 130 may be a device associated with a target domain expert. The determined key elements are presented to the domain expert. The domain expert can identify the main elements to which closed questions should be added. The third computing device 130 sets closed questions in the information gathering sheet 160 according to the domain expert's selection. In this way, the domain expert can use objective data to help design better information gathering sheets, such as better questionnaires.

[0108] Topic models and emotion expression The determination and quantification of target elements have been explained above with reference to Figures 3 to 6B. Target elements can also be extracted from text set 105 using a keyword-assisted topic model. A keyword-assisted topic model (hereinafter referred to as the topic model) connects domain knowledge to the topic model using "anchor words." Anchor words can be used as markers for specific topics; that is, anchor words prompt the topic model to search for topics related to the anchor words. In this way, anchor word-assisted topic models separate different topics from one another. Topic models are useful for identifying topics of interest. Target models include, but are not limited to, Anchored CorEX.

[0109] Therefore, a topic model is used to determine the target elements implicitly included in text set 105. Multiple keywords extracted from text set 105 are used as anchor words in the topic model. In this case, the target element 102 is the target obtained from the target model. The topic model anchors each keyword to a certain topic. Therefore, the topic model is used to determine the correspondence between the target elements and keywords.

[0110] The target elements derived from the topic model are quantified using sentiment representations. Any text analysis method for analyzing sentiment can be used. As an example, a Language Search Word Count (LIWC) dictionary is used. The LIWC dictionary can map words to multiple categories. These categories capture the lexical and semantic features of the text. Categories associated with positive sentiment can be used. LIWC categories that measure positive sentiment are grouped together. Positive sentiment vectors represent the frequency of words belonging to categories of positive sentiment. Estimation of a target element consists of a binarized anchor target variable and a sentiment vector representation.

[0111] Multi-mode model Machine learning models can also be used to analyze information that includes both text and given choices (such as a set of 250 answers). Figure 12A shows a machine learning model 1200 for propensity scores. The language model 1210 within model 1200 is configured to generate feature representations of text 101, which are considered target elements determined from the set of texts 105. Unlike the process described above with reference to Figure 3, the target elements determined in this way are implicit representations.

[0112] The multilayer perception (MLP) layer 1220 is used to generate feature representations of the structured element 103. If the structured element 103 includes both numerical elements (e.g., as shown in Figure 2) and categorical elements (e.g., cabin categories), the MLP layer 1220 includes two MLP layers, one for handling numerical elements and the other for handling categorical elements, respectively.

[0113] The MLP layer 1230 generates features h^e based on the feature representations of text 101 and structured elements 103. Applying the softmax activation function to h^e for feature pairs allows for the determination of propensity scores.

[0114] Model 1200 is trained using a cross-entropy loss function. The language model 1210 can be any suitable type of language feature model, such as a Bidirectional Encoder Representations from Transformers (BERT) model.

[0115] Figure 12B shows a machine learning model 1250 for conditional outcome expectations. The language model 1260 within model 1250 is configured to generate feature representations of text 101, which are considered target elements determined from text set 105. These determined target elements are implicit representations.

[0116] The MLP layer 1270 is used to generate feature representations of the structured element 103. If the structured element 103 includes numerical elements (e.g., as shown in Figure 2) and categorical elements (e.g., cabin categories), the MLP layer 1270 includes two MLP layers for processing the numerical elements and categorical elements, respectively.

[0117] The MLP layer 1280 generates feature h^Q based on the feature representations of text 101 and structured element 103. If result Y is not continuous (e.g., categorical or numerical), a softmax activation function is applied to feature h^Q to determine the expected value of the conditional result. If result Y is continuous, a linear activation function is applied to feature h^Q to determine the expected value of the conditional result.

[0118] If result Y is discontinuous, train model 1250 using the cross-entropy loss function. If result Y is continuous, train model 1250 using the mean-variance (MSE). Similar to language model 1210, language model 1260 may be any suitable type of language feature model, such as the BERT model.

[0119] The key elements are ordered based on either or both propensity scores and / or expected conditional outcomes. This determines the target key elements. Process 700 can also be performed using the machine learning models described in the text.

[0120] Sample device Figure 13 is a block diagram showing a computing device 1300 in which one or more embodiments of the present disclosure are implemented. It should be understood that the computing device 1300 shown in Figure 13 is merely illustrative and should not constitute any limitation on the functionality and scope of the embodiments described herein. The computing device 1300 shown in Figure 13 is used to implement the first computing device 110, the second computing device 120, or the third computing device 130 of Figure 1.

[0121] As shown in Figure 13, the computing device 1300 is in the form of a general-purpose computing device. The components of the computing device 1300 include, but are not limited to, one or more processors or process units 1310, memory 1320, storage device 1330, one or more communication units 1340, one or more input devices 1350, and one or more output devices 1360. The process unit 1310 may be a real processor or a virtual processor and performs various processes according to a program stored in memory 1320. In a multiprocessor system, multiple processing units execute computer executable instructions in parallel to improve the parallel processing capability of the computing device 1300.

[0122] The computing device 1300 generally includes multiple computer memories. Such media include, but are not limited to, volatile and non-volatile media, removable and non-removable media, and any media accessible to the computing device 1300. Memory 1320 may be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or any combination thereof. Storage device 1330 may be removable or non-removable media and may include machine-readable memory such as flash drives, magnetic disks, or any other media accessed within the computing device 1300 that can be used to store information and / or data (e.g., training data for training).

[0123] The computing device 1300 further includes removable / non-removable volatile / non-volatile memory. Not shown in Figure 13, it includes a disk drive for reading from or writing to removable non-volatile disks (e.g., “floppy disks”) and an optical disk drive for reading from or writing to removable non-volatile optical disks. In these cases, each drive may be connected to a bus (not shown) by one or more data medium interfaces. The memory 1320 includes a computer program product 1325 having one or more program modules configured to perform various methods or operations of various embodiments of the present disclosure.

[0124] The communication unit 1340 enables communication with other computing devices via a communication medium. Furthermore, the functions of the components of the computing device 1300 are implemented by a single computing cluster or by multiple computer machines communicating via a communication connection. Thus, the computing device 1300 operates in a network environment using logical connections with one or more other servers, network personal computers (PCs), or other network nodes.

[0125] The input device 1350 is one or more input devices such as a mouse, keyboard, or trackball. The output device 1360 is one or more output devices such as a display, speaker, or printer. The computing device 1300 may also, if necessary, communicate with one or more external devices (not shown) such as a storage device or display device via the communication unit 1340, communicate with one or more devices that enable a user to interact with the computing device 1300, or communicate with any device (e.g., a network card, modem) that enables the computing device 1300 to communicate with one or more other computing devices. Such communication is performed via an input / output (I / O) interface (not shown).

[0126] An exemplary implementation of the present disclosure provides computer-readable memory for storing computer-executable instructions executed by a processor to carry out the methods described above. An exemplary implementation of the present disclosure also provides a computer program product which contains computer-executable instructions that are tangibly stored on non-transient computer-readable memory and executed by a processor to carry out the methods described above.

[0127] Some implementations of this disclosure are shown below.

[0128] In a first aspect, the Disclosure provides a method for processing information. This method includes the steps of extracting a plurality of keywords from an unstructured text set for a target object; grouping at least a portion of the plurality of keywords based on the meaning of the plurality of keywords; grouping at least a portion of the plurality of keywords; and determining a target element corresponding to a group of keywords that represent a particular aspect of the target object, based on the grouping results.

[0129] In some embodiments of the first aspect, the method further includes determining, based on a set of texts, at least one target sentence relating to the target element that reflects a viewpoint on the target element.

[0130] In some embodiments of the first aspect, determining at least one target sentence includes extracting at least one candidate sentence from a text set that contains at least one keyword from a group of keywords, and determining at least one target sentence based on the at least one candidate sentence.

[0131] In some embodiments of the first aspect, at least one target sentence includes at least one sentence that reflects a positive view of the target element and at least one sentence that reflects a negative view of the target element.

[0132] In some embodiments of the first aspect, determining a target element involves determining whether a group of keywords is semantically similar to a target structured element, and if it is determined that the group of keywords is not semantically similar to the structured element, then determining the target element based on the group of keywords.

[0133] In some embodiments of the first aspect, the method further includes determining a measurement of the target element that represents the attention given to the target element, based on a group of keywords and text within a set of texts.

[0134] In some embodiments of the first aspect, determining the measurement includes at least one of determining the frequency of occurrence of a group of keywords in a text and determining the sentiment level of a sentence in the text that contains one of the keywords in the group.

[0135] In some embodiments of the first aspect, the step of extracting multiple keywords includes the step of extracting candidate words from the text of a text set, and the step of extracting candidate words from the text of a text set, and if the number of occurrences of a candidate word in the text set is greater than a threshold number, the candidate word is determined to be one of the multiple keywords.

[0136] In some embodiments of the first aspect, grouping at least some of the keywords includes clustering the keywords to determine a plurality of clusters, each containing at least one keyword, determining the quality of each of the clusters, where the quality represents how semantically the keywords in each cluster are grouped, removing keywords from the plurality of keywords that have a quality lower than a threshold quality to determine the remaining keywords, and grouping the remaining keywords based on the meaning of the remaining keywords.

[0137] In a first aspect, the disclosure provides an electronic device comprising at least one processing circuit. The at least one processing circuit is configured to extract a plurality of keywords from an unstructured text set for a target object, to group at least a portion of the plurality of keywords based on the meaning of the plurality of keywords, to group at least a portion of the plurality of keywords, and to determine a target element corresponding to a group of keywords representing a particular aspect of the target object based on the grouping results.

[0138] In some embodiments of the second aspect, at least one processing circuit is further configured to determine, based on a text set, at least one target sentence relating to the target element that reflects a viewpoint on the target element.

[0139] In some embodiments of the second aspect, determining at least one target sentence includes extracting at least one candidate sentence from a text set that contains at least one keyword from a group of keywords, and determining at least one target sentence based on the at least one candidate sentence.

[0140] In some embodiments of the second aspect, at least one target sentence includes at least one sentence that reflects a positive view of the target element and at least one sentence that reflects a negative view of the target element.

[0141] In some embodiments of the second aspect, determining a target element involves determining whether a group of keywords is semantically similar to a target structured element, and if it is determined that the group of keywords is not semantically similar to the structured element, then determining the target element based on the group of keywords.

[0142] In some embodiments of the second aspect, at least one processing circuit is further configured to determine a measurement of a target element that represents the level of attention given to the target element, based on a group of keywords and text within a set of texts.

[0143] In some embodiments of the second aspect, determining the measurement includes at least one of determining the frequency of occurrence of a group of keywords in the text and determining the sentiment level of sentences in the text that contain the keywords within the group of keywords.

[0144] In some embodiments of the second aspect, the step of extracting multiple keywords includes the step of extracting candidate words from the text of a text set, and the step of extracting candidate words from the text of a text set, and if the number of occurrences of a candidate word in the text set is greater than a threshold number, the candidate word is determined to be one of the multiple keywords.

[0145] In some embodiments of the second aspect, grouping at least some of the keywords of a plurality of keywords includes clustering the plurality of keywords to determine a plurality of clusters, each containing at least one keyword, determining the quality of each of the plurality of clusters, the quality representing how semantically the keywords in each cluster are grouped, and removing keywords from the plurality of keywords that have a quality lower than a threshold quality to determine the remaining keywords.

[0146] In a third aspect, the Disclosure provides computer-readable memory for storing a computer program. The computer program is executed by a processor to implement the information processing method of the first aspect.

[0147] In a fourth aspect, the disclosure provides a method for processing information, which includes the step of obtaining a group of target elements of a target, determined on a set of unstructured text relating to the target, such that each target element represents one aspect of the target; then determining at least one key element for the target based on the group of target elements and a group of structured elements of the target, wherein at least one target element from the group of target elements is different from the group of structured elements.

[0148] In some embodiments of the fourth aspect, determining at least one key element includes selecting a first number of target elements from a group of target elements as part of at least one key element, depending on the degree to which each of the group of target elements has an influence on the target object, and selecting a first number of target elements from the group of target elements, depending on the degree to which each of the group of structured elements has an influence on the target object, and selecting a second number of structured elements from the group of structured elements as part of at least one key element.

[0149] In some embodiments of the fourth aspect, determining at least one key element involves selecting a third number of elements from a group of target elements and a group of structured elements as at least one key element, depending on the degree to which each of a group of target elements has an influence on the target object and the degree to which each of a group of structured elements has an influence on the target object.

[0150] In some embodiments of the fourth aspect, determining key elements includes determining a first measure for each of a group of target elements representing the attention given to the corresponding target element by analyzing the sentiment of the text in the text set, and determining a second measure for each of a group of structured elements representing the attention given to the corresponding structured element, wherein the second measure is consistent with the first measure in terms of the measurement scale, and determining at least one key element based on the first and second measures.

[0151] In some embodiments of the fourth aspect, the first measurement is determined based on the number of occurrences of the first keyword in the first sentence for a target element, which includes determining from the text of a text set a first sentence containing a first keyword corresponding to the target element for a target element of a group of target elements, which includes determining from the text of a text set a first sentence containing a first keyword corresponding to the target element, which includes determining for a target element of a group of target elements, which includes determining from the text a second sentence containing a second keyword corresponding to the structured element for a structured element of a group of structured elements, which includes determining a second sentence containing a second keyword corresponding to the structured element for a structured element, which includes determining for a structured element of a group of structured elements, which includes determining a second measurement based on the number of occurrences of the second keyword in the second sentence.

[0152] In some embodiments of the fourth aspect, the step of determining a first measurement includes determining, for a target element of a group of target elements, a first measurement of the target element based on the sentiment level of the sentence, including the step of determining, for the target element, a sentence that includes the keyword corresponding to the target element and has sentiment, from the text of a text set, and the step of determining, for the target element of a group of target elements, a first measurement of the target element based on the sentiment level of the sentence, and the step of determining, for the target element of a group of target elements, a second measurement of the structured element based on the structured element of a group of structured elements and answers to closed questions about the structured element.

[0153] In some embodiments of the fourth aspect, determining at least one key element based on a first measurement and a second measurement includes determining the degree to which each of a group of target elements and a group of structured elements has an influence on the target object based on the first measurement and the second measurement, and determining the degree of influence on the target object based on the first measurement and the second measurement, and selecting at least one key element from the group of target elements and the group of structured elements based on the degree of influence.

[0154] In some embodiments of the fourth aspect, the degree of influence is determined based on at least one of linear regression, logical regression, and Sharley value.

[0155] In a fifth aspect, the disclosure provides an electronic device comprising at least one processing circuit, the at least one processing circuit configured to acquire a group of target elements of a target object, each of which is determined based on an unstructured text set relating to the target object, and each target element represents one aspect of the target object, and to determine at least one key element for the target object based on the group of target elements and a group of structured elements of the target object, wherein at least one of the target elements in the group is different from the group of structured elements.

[0156] In some embodiments of the fifth aspect, at least one processing circuit is further configured to select a first number of target elements from a group of target elements as part of at least one key element, depending on the degree to which each of the group of target elements has an influence on the target object, and to select a second number of structured elements from a group of structured elements as part of at least one key element, depending on the degree to which each of the group of structured elements has an influence on the target object.

[0157] In some embodiments of the fifth aspect, at least one processing circuit is further configured to select a third number of elements as at least one key element from a set of the sum of a group of target elements and a group of structured elements, depending on the degree to which each of a group of target elements has an influence on the target object and the degree to which each of a group of structured elements has an influence on the target object.

[0158] In some embodiments of the fifth aspect, at least one processing circuit is further configured to determine a first measurement for each of a group of target elements representing the attention given to the corresponding target element by analyzing the sentiment of the text in a text set, the first measurement then determines a second measurement for each of a group of structured elements representing the attention given to the corresponding target element by analyzing the sentiment of the text in a text set, the second measurement being consistent with the first measurement in terms of measurement scale, and at least one key element is determined based on the first and second measurements.

[0159] In some embodiments of the fifth aspect, at least one processing circuit is further configured to determine, for a target element in a group of target elements, a first measurement of the target element based on the number of occurrences of the first keyword in the first sentence, which includes a first keyword corresponding to the target element and has sentiment, for a structured element in a group of structured elements, a second measurement of the structured element is determined based on the number of occurrences of the second keyword in the second sentence, which includes determining, for a structured element in a group of structured elements, a second sentence containing a second keyword corresponding to the structured element and having sentiment, for a structured element in a group of structured elements, a second measurement of the structured element is determined based on the number of occurrences of the second keyword in the second sentence.

[0160] In some embodiments of the fifth aspect, at least one processing circuit determines a first measurement of a target element based on the sentiment level of a sentence, for a target element of a group of target elements, and for a structured element of a group of structured elements, it determines a second measurement of a structured element based on answers to closed questions about the structured element.

[0161] In some embodiments of the fifth aspect, at least one processing circuit is further configured to determine the degree of influence of a group of target elements and a group of structuring elements on a target object based on a first measurement and a second measurement, and selects at least one key element from the group of target elements and the group of structuring elements based on the degree of influence.

[0162] In some embodiments of the fifth aspect, the degree of influence is determined based on at least one of linear regression, logical regression, and Sharley value.

[0163] In a sixth aspect, the disclosure provides computer-readable memory for storing a computer program. The computer program is executed by a processor to implement the method of the fourth aspect.

[0164] In a seventh aspect, the disclosure provides an information processing method, which includes the step of obtaining a group of target elements of a target, which are determined based on an unstructured text set relating to the target, and each target element represents one aspect of the target. The method then presents an information gathering sheet for collecting a description of the target based on at least one key element of the target, which is determined from a group of structured elements and a group of target elements of the target.

[0165] In some embodiments of the seventh aspect, the text of the text set is obtained from answers to open questions in an information gathering sheet, the information gathering sheet includes corresponding closed questions relating to a group of structured elements, and presenting the information gathering sheet includes presenting an updated version of the information gathering sheet, which includes updated closed questions based on at least one key element.

[0166] In some embodiments of the seventh aspect, presenting an updated version of the information gathering sheet includes presenting an updated version that includes a first closed question relating to a first target element if it is determined that at least one key element includes a first target element of a group of target elements.

[0167] In some embodiments of the seventh aspect, presenting an updated version of the information gathering sheet includes presenting an updated version in which the second closed question relating to the first structured element is removed if it is determined that at least one element does not contain the first structured element of a group of structured elements.

[0168] In some embodiments of the seventh aspect, at least one key element includes a second target element of a group of target elements, and the method further includes detecting answers to open questions in an information gathering sheet while the information gathering sheet is being presented, and, in response to detecting that answers have been provided, presenting a hint describing the second target element.

[0169] In some embodiments of the seventh aspect, presenting an information collection sheet includes presenting at least one key element, detecting the selection of at least one key element while the at least one key element is presented, detecting the selection of at least one key element, and presenting an information collection sheet containing closed questions relating to the selected key element in response to the detection of the selection of at least one key element.

[0170] In an eighth aspect, the disclosure provides an electronic device comprising at least one processing circuit, the at least one processing circuit configured to acquire a group of target elements of a target, where the group of target elements is determined based on an unstructured text set relating to a target, and each target element represents one aspect of the target, and presents an information gathering sheet for collecting a description of the target based on at least one key element of the target, where the at least one key element is determined from a group of structured elements and a group of target elements of the target.

[0171] In some embodiments of the eighth aspect, the text of the text set is obtained from answers to open questions in an information gathering sheet, the information gathering sheet includes corresponding closed questions relating to a group of structured elements, and at least one processing circuit is further configured to present an updated version of the information gathering sheet, including updated closed questions based on at least one key element.

[0172] In some embodiments of the eighth aspect, at least one processing circuit is further configured to present an updated version including a first closed question relating to the first target element if it is determined that at least one key element includes a first target element of a group of target elements.

[0173] In some embodiments of the eighth aspect, at least one processing circuit is further configured to present an updated version in which the second closed question relating to the first structured element is removed if it is determined that at least one element does not contain a first structured element of a group of structured elements.

[0174] In some embodiments of the eighth aspect, the at least one key element includes a second target element of a group of target elements, and the at least one processing circuit is configured to detect answers to open questions in the information gathering sheet while the information gathering sheet is being presented, and in response to detecting that answers have been provided, presents a hint describing the second target element.

[0175] In some embodiments of the eighth aspect, the at least one processing circuit is further configured to present the at least one key element, the at least one processing circuit to detect the selection of the at least one key element while the at least one key element is being presented, and in response to the detection of the selection of the at least one key element, to present an information gathering sheet including closed questions relating to the selected key element.

[0176] In the ninth aspect, a computer-readable memory is provided for storing a computer program. The computer program is executed by a processor to implement the method of the seventh aspect.

[0177] Various aspects of this disclosure are described herein with reference to flowcharts and / or block diagrams of methods, devices, equipment, and computer program products implemented in accordance with this disclosure. It should be understood that each block in the flowcharts and / or block diagrams, and each combination of blocks in the flowcharts and / or block diagrams, may be implemented by computer-readable program instructions.

[0178] These computer-readable program instructions are provided to a processing unit of a general-purpose computer, a dedicated computer, or other programmable data processing device to manufacture a machine that, when these instructions are executed via the processing unit of the computer or other programmable data processing device, generates means to implement functions / operations defined in one or more blocks in a flowchart and / or block diagram. These computer-readable program instructions may also be stored in computer-readable memory that operates the computer, programmable data processing device, and / or other device in a particular way, thereby including a manufactured article that implements various aspects of functions / operations defined in one or more blocks in a flowchart and / or block diagram.

[0179] Computer-readable program instructions are loaded into a computer, another programmable data processing device, or other device to generate a computer implementation process in which a series of operational steps are performed on the computer, another programmable data processing device, or other device, so that instructions executed on the computer, another programmable data processing device, or other device perform functions / operations defined in one or more blocks in a flowchart and / or block diagram.

[0180] The attached flowcharts and block diagrams illustrate the architecture, functionality, and operation of possible embodiments of the system, method, and computer program product according to several embodiments of this disclosure. In this regard, each block in the flowchart or block diagram represents a module, program segment, or part of an instruction containing one or more executable instructions for implementing a given logical function. In some alternative implementations, the functions shown in the blocks may occur in a different order than shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may be executed in reverse order depending on the functions involved. Each block in the block diagram and / or flowchart, and / or combinations of blocks in the block diagram and / or flowchart, may be implemented in a dedicated hardware-based system that performs a given function or operation, or in a combination of dedicated hardware and computer instructions.

[0181] While embodiments of the present disclosure have been described above, the above description is illustrative, not exhaustive, and not limited to the embodiments disclosed. Many modifications and changes will be obvious to those skilled in the art without departing from the scope and spirit of each embodiment described. The choice of terms used herein is intended to best describe each embodiment disclosed herein, the principle of each embodiment disclosed herein, its practical application, or improvements in the technology in the market, or to make them understandable to other ordinary people of the art.

Claims

1. Methods of information processing using computing devices, including the following: The computing device extracts multiple keywords from an unstructured text set that is related to a target and contains multiple sentences. The computing device groups at least some of the multiple keywords based on their meanings and generates a group containing one or more keywords. The computing device determines a cluster center consisting of the group of keywords, and determines the keyword with the semantic characteristics closest to the cluster center among the keywords of the group as the target element. and The computing device extracts candidate sentences containing the target element from the sentences of the unstructured text set, and uses the extracted candidate sentences as at least one target sentence.

2. The method according to claim 1, wherein the at least one target sentence includes at least one sentence that reflects a positive view of the target element and a sentence that reflects a negative view of the target element.

3. The method according to claim 1, further comprising: The computing device determines the frequency of occurrence of a group of keywords in the text of the unstructured text set, and determines the level of attention given to the determined target element based on the frequency of occurrence.

4. Extracting the aforementioned plurality of keywords includes the following according to the method of claim 1: The computing device extracts candidate words from the text of the unstructured text set; and The computing device determines that a candidate word is one of several keywords if the number of occurrences of the candidate word in the unstructured text set is greater than a threshold number.

5. The method according to claim 1, wherein grouping at least some of the keywords mentioned above includes the following: The computing device clusters the multiple keywords to identify multiple clusters, and each cluster consists of at least one keyword. The computing device determines the quality of each of the multiple clusters based on the sum of the squared distances of keywords within each cluster or the contour coefficient, and the quality represents the degree to which the keywords within each cluster are semantically grouped. The computing device determines the remaining keywords by removing keywords in the cluster that have a quality lower than the threshold quality from among multiple keywords, and The remaining keywords will be grouped based on their meaning.

6. It includes at least one processing circuit, and the at least one processing circuit is Extract multiple keywords from an unstructured text set that is relevant to the target audience and contains multiple sentences. Based on the meanings of the aforementioned multiple keywords, at least a portion of the aforementioned multiple keywords are grouped together to generate a group containing one or more keywords. A cluster center consisting of the keywords of the aforementioned group is determined, and the keyword with the semantic characteristics closest to the cluster center among the keywords of the aforementioned group is determined as the target element. An electronic device configured to extract candidate sentences containing the target element from the sentences of the unstructured text set, and to use the extracted candidate sentences as at least one target sentence.

7. The electronic device according to claim 6, wherein the at least one target sentence includes at least one sentence reflecting a positive view of the target element and a sentence reflecting a negative view of the target element.

8. The at least one processing circuit is, The electronic device according to claim 6, configured to determine the frequency of occurrence of a group of keywords in the text of the unstructured text set, and to represent the degree of attention of the target element determined based on the frequency of occurrence.

9. The electronic device according to claim 6, wherein the at least one processing circuit is configured to include the following: Extracting candidate words from the text of the aforementioned unstructured text set; and If the number of occurrences of a candidate word in the aforementioned unstructured text set is greater than a threshold number, the candidate word is determined to be one of several keywords.

10. The electronic device according to claim 6, wherein the at least one processing circuit is configured to include: The aforementioned multiple keywords are clustered to identify multiple clusters, and each cluster consists of at least one keyword. The quality of each cluster is determined based on the sum of the squared distances of keywords within each cluster or the contour coefficient. This quality represents the degree to which the keywords within each cluster are semantically grouped together. To determine the remaining keywords, remove keywords in clusters that have a quality lower than the threshold quality from among multiple keywords, and The remaining keywords will be grouped based on their meaning.

11. A computer-readable memory that stores a computer program executable by a processor for carrying out an information processing method, The aforementioned information processing method includes the following: Extracting multiple keywords from an unstructured text set that includes multiple sentences related to the target audience; Based on the meanings of the aforementioned multiple keywords, group at least some of the aforementioned multiple keywords to generate a group containing one or more keywords, and A cluster center consisting of the keywords from the aforementioned group is determined, and the keyword with the semantic characteristics closest to the cluster center among the keywords from the aforementioned group is determined as the target element. Extract candidate sentences containing the target element from the sentences of the unstructured text set, and use the extracted candidate sentences as at least one target sentence.