Artificial intelligence brand market share analysis method, system, storage medium and program product
Patent Information
- Application Number
- CN202611105165.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-24
- Publication Date
- 2026-08-21
AI Technical Summary
如果仅按照原始字符串进行统计,容易出现同一品牌被拆分统计、竞品被遗漏统计、无关实体被误计入品牌提及等问题
(1)本发明以人工智能回答文本作为分析对象,根据品牌在回答中的提及情况、回答覆盖情况、问题覆盖情况、出现位置和情绪倾向计算AI品牌市场份额,相比仅依赖销售额、搜索量或社交声量的传统统计方式,更适用于生成式人工智能问答场景,能够反映品牌在人工智能回答中的可见度、推荐强度和竞争位置。
Smart Images

Figure CN122617451A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis technology, and in particular to AI-based brand market share analysis methods, systems, storage media, and program products. Background Technology
[0002] As generative AI question-answering platforms gradually become an important entry point for users to understand product services, brand information, and purchase advice, whether a brand is mentioned in AI answers, whether it is at the top of the recommendation list, and whether it appears at the same time as competitors have gradually become important reference indicators for brand digital operations.
[0003] Traditional brand market share metrics, typically based on sales volume, search volume, media buzz, or social media discussion, reflect a brand's performance in transactional, search, or communication scenarios. However, they struggle to capture a brand's visibility, recommendation strength, and competitive position in AI-powered question-and-answer scenarios. Therefore, quantifying brand performance in AI-generated answers has become a crucial issue in brand data analytics.
[0004] In existing technologies, some solutions revolve around optimizing generative engines, primarily focusing on how to optimize the output content or retrieved documents of generative engines to improve the presentation of content in generated answers. Other solutions focus on cross-platform competitor reputation analysis, typically using historical reputation texts from social media platforms, media platforms, or other public platforms as the analysis object to statistically analyze the performance of different brands in terms of reputation sentiment, dissemination popularity, or platform voice. These technologies can help brands understand external content dissemination or user reviews to some extent, but their analysis objects and statistical methods are mainly focused on web page content, social text, or reputation data, rather than specifically targeting AI-generated answer text.
[0005] However, on the one hand, in AI-generated answering scenarios, the same answer may contain multiple brands simultaneously, and the same brand may appear in various forms such as Chinese name, English name, abbreviation, alias, product name, or corporate entity. If statistics are only based on the raw strings, problems such as the same brand being split into separate statistics, competitors being omitted from the statistics, and irrelevant entities being mistakenly included in brand mentions can easily occur. On the other hand, different questions have varying degrees of impact on brand decisions, and the business value of a brand appears at the top of the recommendation list, in the first paragraph of the text, or at the end of the answer is not the same; simply summing the number of mentions is insufficient to accurately reflect the brand's true competitive position in AI-generated answers. Furthermore, existing technologies lack unified rules for handling anomalies such as low-confidence entities, duplicate answers, brand ambiguity, and product name duplication, and it is also difficult to form multi-layered comparative analyses of the product and its competitors across dimensions such as question, platform, cycle, sentiment, and position.
[0006] Therefore, how to provide an AI brand market share analysis technology for AI-driven answering scenarios, which combines question value, location of appearance, sentiment, and multi-brand comparison relationships to accurately calculate the brand market share of the product and its competitors in AI-driven answers, is a problem that urgently needs to be solved. Summary of the Invention
[0007] This invention provides an AI-based brand market share analysis method, system, storage medium, and program product to address the aforementioned problems in the prior art.
[0008] According to a first aspect of the present invention, an AI-based brand market share analysis method is provided.
[0009] In one embodiment, the AI brand market share analysis method includes: generating an analysis task based on user configuration, including configurations for the product itself, competitors, brand entities, and analysis scope; reading AI answer records according to the analysis task to determine the set of answers to be analyzed; performing brand entity recognition on the answer text in the set of answers to be analyzed to obtain candidate brand entities; the candidate brand entities include entity name, text position, context fragment, and recognition confidence; performing standard brand merging on the candidate brand entities based on the brand entity configuration to obtain valid brand mentions bound to standard brand codes; generating classification records based on the question text to which the valid brand mentions belong, and determining question weights based on the classification records; determining the occurrence position, recommendation position, and position weight by the text position of the valid brand mentions in the answer text, and identifying the sentiment tendency corresponding to the valid brand mentions; calculating the AI brand market share of the product itself and competitors based on the valid brand mentions, question weights, position weights, and sentiment tendency; and generating multi-brand comparative analysis results based on the AI brand market share.
[0010] According to a second aspect of the present invention, an AI-based brand market share analysis system is provided.
[0011] In one embodiment, the AI brand market share analysis system includes: a brand entity configuration module, used to generate an analysis task based on user configuration, including the product itself, competitors, brand entity configuration, and analysis scope configuration; a question scope configuration module, used to read AI answer records according to the analysis task and determine the set of answers to be analyzed; a brand entity recognition module, used to perform brand entity recognition on the answer text in the set of answers to be analyzed to obtain brand candidate entities; the brand candidate entities include entity name, text position, context fragment, and recognition confidence; a standard brand merging module, used to perform standard brand merging on the brand candidate entities based on the brand entity configuration to obtain valid brand mentions bound to standard brand codes; a question layering module, used to generate classification records based on the question text to which the valid brand mentions belong, and determine the question weights based on the classification records; determine the occurrence position, recommendation position, and position weight by the text position of the valid brand mentions in the answer text, and identify the sentiment tendency corresponding to the valid brand mentions; a market share calculation module, used to calculate the AI brand market share of the product itself and competitors based on the valid brand mentions, question weights, position weights, and sentiment tendencies; and a comparative analysis module, used to generate multi-brand comparative analysis results based on the AI brand market share.
[0012] According to a third aspect of the present invention, a storage medium is provided.
[0013] In one embodiment, the storage medium is a computer-readable storage medium on which a computer program is stored, which, when executed by a processor, implements the steps of the above method.
[0014] According to a fourth aspect of the present invention, a computer program product is provided.
[0015] In one embodiment, the computer program product includes a computer program that, when executed by a processor, implements the steps of the method described above.
[0016] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: (1) This invention uses AI-generated answer text as the analysis object. It calculates the market share of AI brands based on the brand's mention in the answer, answer coverage, question coverage, appearance location and sentiment tendency. Compared with traditional statistical methods that rely solely on sales, search volume or social voice, it is more suitable for generative AI question-and-answer scenarios and can reflect the brand's visibility, recommendation strength and competitive position in AI answers.
[0017] (2) This invention, through brand entity configuration and standard brand merging mechanism, merges the Chinese name, English name, abbreviation, alias, product name, official website domain name and corporate entity of the same brand into a unified standard brand code, which can reduce the problems of duplicate counting, omission and miscounting caused by multiple brand names, and make the mention statistics of this product and competitors have the same calculation caliber.
[0018] (3) In the process of brand entity recognition, the present invention integrates dictionary matching score, context rule score, entity recognition model probability and semantic disambiguation score, and handles abnormal entities by determining confidence threshold and reducing the weight of low confidence entity proportion. This can reduce the impact of irrelevant entities, ambiguous entities and product name duplication on the share calculation results, and improve the accuracy and stability of brand recognition results in artificial intelligence answering scenarios.
[0019] (4) This invention can generate report results by combining problem stratification, position weight, sentiment tendency and multi-brand comparison analysis, so that brand owners can not only see the overall market share difference between their own products and competitors, but also further identify the advantages and disadvantages under different problems, different platforms, different time periods and different sentiments, thereby providing a basis for brand content construction, problem monitoring and competitor strategy adjustment.
[0020] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description
[0021] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0022] Figure 1 This is a flowchart illustrating an AI brand market share analysis method according to an exemplary embodiment; Figure 2 This is a specific implementation diagram illustrating an AI brand market share analysis method according to an exemplary embodiment; Figure 3 This is a schematic diagram illustrating the market share analysis of AI-based brand entities and the merging relationship of standard brands, based on an exemplary embodiment. Figure 4 This is a schematic diagram of a comprehensive AI brand market share calculation model according to an exemplary embodiment; Figure 5 This is a schematic diagram of the module structure of an AI brand market share analysis system according to an exemplary embodiment; Figure 6 This is a schematic diagram illustrating the principle of an AI brand market share analysis system according to an exemplary embodiment; Figure 7This is a schematic diagram of the structure of a computer device according to an exemplary embodiment. Detailed Implementation
[0023] The following description and accompanying drawings fully illustrate specific embodiments described herein to enable those skilled in the art to practice them. Some portions and features of certain embodiments may be included in or replace portions and features of other embodiments. The scope of the embodiments herein includes the entire scope of the claims and all available equivalents thereof. The various embodiments described herein are presented in a progressive manner, with each embodiment focusing on its differences from other embodiments; similar or identical parts between embodiments can be referred to interchangeably.
[0024] The modules in the apparatus or system of this application can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0025] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0026] Figure 1 An embodiment of an AI-based brand market share analysis method of the present invention is shown.
[0027] In this optional embodiment, the AI brand market share analysis method includes: Step S101, generating an analysis task based on user configuration, including configurations for the product itself, competitors, brand entities, and analysis scope; Step S102, reading AI answer records according to the analysis task to determine the set of answers to be analyzed; Step S103, performing brand entity recognition on the answer text in the set of answers to be analyzed to obtain candidate brand entities; the candidate brand entities include entity name, text position, context fragment, and recognition confidence; Step S104, performing standard brand merging on the candidate brand entities based on the brand entity configuration to obtain effective brand mentions bound to standard brand codes; Step S105, generating classification records based on the question text to which the effective brand mentions belong, and determining the question weights based on the classification records; determining the occurrence position, recommendation position, and position weight through the text position of the effective brand mentions in the answer text, and identifying the sentiment tendency corresponding to the effective brand mentions; Step S106, calculating the AI brand market share of the product itself and competitors based on the effective brand mentions, question weights, position weights, and sentiment tendencies; Step S107, generating multi-brand comparative analysis results based on the AI brand market share.
[0028] (I) Overall Approach: It should be noted that this invention uses user-configured information such as the product itself, competitors' products, brand aliases, product names, official website domains, project scope, question scope, AI platform scope, and time period as input for the analysis task, and uses the system's AI answer records, subscription tracking results, or user-imported question and answer files within the user's authorized scope as the analysis objects.
[0029] Specifically, brand entity recognition is first performed on the answer text in the set of answers to be analyzed to obtain brand candidate entities, text positions, context fragments and recognition confidence; then, based on brand main body configuration, alias dictionary, official website main body, product affiliation relationship and exclusion word rules, the candidate entities are merged into standard brands.
[0030] Specifically, after brand consolidation is completed, the position, recommendation ranking, question category, AI platform, time period, and sentiment of the standard brand in each answer are identified. Based on the number of mentions, answer coverage, question coverage, position weight, sentiment weight, and platform weight, the AI brand market share of the product and its competitors is calculated, ultimately generating a multi-brand comparative analysis result of the product and its competitors.
[0031] (ii) System input data: The system input data of this invention includes, but is not limited to, the following fields: product name, competitor name, brand standard name, brand Chinese name, brand English name, abbreviation, historical name, product name, series name, official website domain name, corporate entity, commonly misspelled names, excluded words, industry category, project scope, problem scope, AI platform scope, time period, AI answer records in the system, subscription tracking results, and user-imported question and answer files.
[0032] The AI answer record can include the question text, answer text, AI platform, model name, time, project, brand configuration, and existing parsing results; the user-imported question and answer file can include fields such as question, answer, platform, time, and source project. All of the above inputs are limited to the scope authorized by the user.
[0033] It should be added that, to address the scenario where users may not have authorized access to the system's AI answer records, leading to missing input data, this invention provides a three-tiered data acquisition mechanism: The first tier consists of user-authorized access to the system's AI answer records and subscription tracking results, with this being the preferred data source; the second tier consists of user-imported question-and-answer files, supporting Excel, CSV, and JSON formats, with fields including question text, answer text, AI platform name, answer time, and source project identifier; the third tier consists of the system's built-in public AI question-and-answer sample library. When the above two tiers of data are unavailable, the system matches the most relevant question-and-answer samples from the sample library based on the brand and question scope configured by the user as the analysis input, and clearly marks the data source and quality level in the report. This three-tiered data acquisition mechanism ensures that the system can execute the analysis process normally under any data availability conditions.
[0034] Specifically, the three-level data acquisition mechanism includes: the system uses the brand name, competitor name, industry category and question scope configured by the user as matching input to calculate the relevance score for each question and answer record in the publicly available AI question and answer sample library. The relevance calculation adopts a multi-dimensional weighted matching method: (Dimension 1) Semantic similarity calculation between the question text and the user-configured question range, using the Sentence-BERT model to output cosine similarity, with a weight of 0.40; (Dimension 2) Whether the answer text contains an exact or fuzzy match of the user-configured brand name or competitor name, with a score of 1.00 for a match and 0.00 for a miss, with a weight of 0.30; (Dimension 3) The degree of matching between the industry category tags of the question and answer record and the industry category configured by the user, with a score of 1.00 for a perfect match, 0.60 for the same parent category, and 0.00 for irrelevant categories, with a weight of 0.20; (Dimension 4) The timeliness of the question and answer record, with a score of 1.00 for records within 30 days, 0.70 for records within 90 days, and 0.40 for records older than 90 days, with a weight of 0.10. The final relevance score is calculated as Relevance = Σ(dimension score × dimension weight). The system sorts the data in descending order of Relevance and takes the top-N items with a default N=100 as the analysis input. The report also marks the sample database matching data and the corresponding Relevance score range.
[0035] In this optional embodiment, the analysis task is generated based on the set of brands to be followed, the brand entity configuration, and the analysis scope configuration. The set of brands to be followed includes the user-configured products and competitors. The brand entity configuration includes the brand alias, product name, official website domain name, corporate entity, and excluded keywords corresponding to the set of brands to be followed. The analysis scope configuration includes the user-configured project scope, question scope, AI platform scope, time period, and industry category. Among them, the brand entity configuration is used for standard brand merging, the question scope, AI platform scope, and time period are used to read AI answer records, and the industry category is used to match sample questions and answers from the public AI question and answer sample library.
[0036] In this optional embodiment, the determination of the set of answers to be analyzed is performed through a three-level data acquisition mechanism, specifically including: based on the analysis task, reading a first candidate set of answers from the system's AI answer records and subscription tracking results within the user's authorized scope; when the first candidate set of answers is unavailable, reading a second candidate set of answers from the user-imported question and answer file; when both the first and second candidate set of answers are unavailable, matching sample questions and answers from a public AI question and answer sample library based on the product, competitors, industry category, and question scope to obtain a third candidate set of answers; and determining the set of answers to be analyzed from the first, second, or third candidate set of answers according to the currently available data source, and configuring a data source identifier and quality level for the set of answers to be analyzed.
[0037] In this optional embodiment, sample questions and answers are matched from a publicly available AI question-and-answer sample library to obtain a third candidate answer set, including: obtaining a question relevance score based on the semantic similarity between the question scope and the sample question text; generating a brand hit score based on whether the sample answer text matches the product or a competitor's product; determining an industry matching score based on the degree of matching between the sample industry tags and the industry category corresponding to the analysis task; forming a timeliness score based on the proximity between the sample answer time and the time period; weighting and fusing the question relevance score, brand hit score, industry matching score, and timeliness score to obtain the sample relevance; sorting the sample questions and answers in the publicly available AI question-and-answer sample library according to the sample relevance, and selecting a preset number of sample questions and answers at the top of the sorted list as the third candidate answer set.
[0038] In this optional embodiment, brand entity recognition is performed on the answer text in the set of answers to be analyzed to obtain brand candidate entities, including: preprocessing the answer text to obtain a sentence sequence; performing multi-source dictionary matching on the sentence sequence based on a brand recognition dictionary system, and performing context rule matching based on context fragments to obtain dictionary matching scores and context rule scores; performing sequence labeling on the sentence sequence through a brand entity recognition model to obtain entity recognition model probabilities; disambiguating the initial candidate entities within a preset confidence interval to obtain disambiguation scores; determining the recognition confidence level based on the dictionary matching score, context rule score, entity recognition model probability, and disambiguation score, and outputting brand candidate entities.
[0039] In this optional embodiment, the brand identification dictionary system includes a precise brand dictionary, a fuzzy alias dictionary, and an exclusion word dictionary. The precise brand dictionary is constructed based on the brand's standard name, English name, abbreviation, and product name, and is used for precise matching of sentence sequences. The fuzzy alias dictionary is constructed based on common misspelled names, colloquial terms, and abbreviation variations, and is used for fuzzy matching of text fragments that are not precisely matched. The exclusion word dictionary is constructed based on place names, common nouns, industry terms, and user-defined exclusion words, and is used to filter text fragments that should not be identified as brand entities. When there are word conflicts between the precise brand dictionary, the fuzzy alias dictionary, and the exclusion word dictionary, the corresponding text fragment is filtered first based on the exclusion word dictionary.
[0040] In this optional embodiment, the identification confidence score is used for threshold determination and low-confidence entity processing of brand candidate entities. Specifically, it includes: setting several candidate confidence thresholds based on labeled AI answer samples; determining whether a brand candidate entity is a brand entity or a non-brand entity according to each candidate confidence threshold, and calculating the precision and recall corresponding to each candidate confidence threshold; calculating the harmonic evaluation value corresponding to each candidate confidence threshold based on the precision and recall, and determining the candidate confidence threshold corresponding to the maximum value of the harmonic evaluation value as the preset inclusion threshold; when the identification confidence score of a brand candidate entity is greater than or equal to the preset inclusion threshold, the corresponding brand candidate entity is determined as a threshold-passing candidate entity; when the identification confidence score of a brand candidate entity is lower than the preset inclusion threshold, the corresponding brand candidate entity is determined as a low-confidence candidate entity; and proportionally reducing the statistical weight of the low-confidence candidate entity according to the ratio between the identification confidence score and the preset inclusion threshold, or marking the low-confidence candidate entity as pending confirmation.
[0041] (III) Brand Entity Identification Mechanism: Specifically, the system performs brand entity recognition on the response text. The recognition objects include brand name, product name, company name, English name, abbreviation, and expressions in the context that can point to the brand entity.
[0042] The brand entity recognition in this invention employs the following multi-level method steps: Step 1: Text Preprocessing. The AI response text is segmented into sentences, words, and cleaned of special characters. The response text is split into sentence sequences according to natural paragraphs and punctuation marks, and HTML tags and irrelevant formatting characters are removed.
[0043] Step 2: Multi-source dictionary matching. A brand identification dictionary system is constructed, comprising three layers: ① A precise brand dictionary, storing the brand's standard name / English name / abbreviation / product name, using the Aho-Corasick multi-pattern matching algorithm of the AC automaton to achieve synchronous scanning with low time complexity; ② A fuzzy alias dictionary, storing common misspellings and colloquial terms, using edit distance ≤ 2 and pinyin similarity of double-pinyin encoding for fuzzy matching; ③ An exclusion word dictionary, storing exclusion words such as place names, common nouns, and industry terms that should not be identified as brand entities, filtering them during the matching stage.
[0044] Specifically, the construction method of the three-layer dictionary system includes: (I) Steps for constructing a precise brand dictionary: Step 1: Export the standard Chinese name, English name, abbreviation list and product name list of all brands in batches from the system brand configuration table to form an initial term set; Step 2: Standardize each term, including unifying uppercase and lowercase, removing redundant spaces and special characters, and unifying full-width and half-width characters; Step 3: Group the standardized terms by brand ID and construct the AC automaton Aho-Corasick data structure, specifically: construct a Trie tree in BFS order, calculate the failure link of each node, preprocess the output link, and realize multi-mode synchronous scanning; Step 4: Periodically incrementally update, and rebuild the AC automaton when the brand configuration table adds or modifies terms. (II) Steps for constructing a fuzzy alias dictionary: ① Collect common misspellings, colloquial names, and abbreviation variations of brands, including unmatched query terms from user history input logs and manually compiled alias tables; ② Calculate the Levenshtein Distance between each alias and the standard brand name, retaining entries with an edit distance ≤ 2; ③ Calculate the pinyin similarity for Chinese aliases, using a double-pinyin encoding scheme to convert Chinese characters into initial consonant + final vowel encoding sequences, calculate the Jaccard similarity coefficient of the encoding sequences, and include entries with a threshold ≥ 0.70 in the dictionary; ④ Establish a mapping index from aliases to standard brand IDs. (III) Exclusion Term Dictionary Construction Steps: Step 1: Extract basic exclusion terms from a general place name database containing provinces, cities, counties, and districts; a database of the top 5000 most common names; and an industry terminology database containing common conceptual terms for various industries. Step 2: Load the corresponding industry-specific exclusion term table according to the industry category configured by the user. Step 3: Incorporate user-defined exclusion terms through the exclusion term field in the brand configuration table. Step 4: Build a hash index for the exclusion terms and perform fast filtering in terms of time complexity during the dictionary matching stage of brand entity recognition. The three-layer dictionary automatically performs a full consistency check every 24 hours to ensure no conflicts between dictionaries, such that a single term cannot exist simultaneously in both the precise brand dictionary and the exclusion term dictionary.
[0045] Step 3: Contextual Rule Matching. Candidate entities matched to the dictionary undergo secondary validation based on predefined rule templates: ① Increase confidence by 0.1 to 0.2 when brand indicator words such as "brand," "recommendation," or "ranking" appear before or after the entity; ② Decrease confidence by 0.15 to 0.3 when a candidate entity has a place name suffix before or after it but no brand indicator words; ③ Increase confidence by 0.15 when a candidate entity and other confirmed brand entities in the response are in the same enumeration structure.
[0046] Step 4: Entity Recognition Model Inference. The Brand-NER entity recognition model, fine-tuned based on a pre-trained language model (BERT-wwm-ext or RoBERTa), is used to perform sequence labeling on the response text. The labeling categories are B-BRAND, I-BRAND, B-PRODUCT, I-PRODUCT, and O, outputting the entity category probability distribution for each token. The Brand-NER model uses BERT-wwm-ext or RoBERTa as the encoder, with a fully connected linear layer and a Conditional Random Field (CRF) layer connected sequentially at the top layer. The input text is segmented into word embedding vector sequences using WordPiece. A 12-layer Transformer encoder extracts contextual semantic features. The linear layer maps the hidden state of each token to a 5-dimensional label score. The CRF layer globally decodes the overall sequence labeling results based on the label transition matrix, outputting the optimal sequence of labels for the five categories: B-BRAND, I-BRAND, B-PRODUCT, I-PRODUCT, and O.
[0047] Specifically, the training process of the brand entity recognition model is as follows: First, pre-training is performed on the CLUE Chinese corpus to obtain the basic weights of BERT-wwm-ext or RoBERTa; then, an AI response sample containing more than 5000 manually annotated entries is constructed as a fine-tuning dataset, with brand entities in each sample independently annotated by at least two annotators and consistent results are taken; the AdamW optimizer is used, with a learning rate of 2e-5, a warm-up ratio of 0.1, a batch size of 32, a maximum sequence length of 512, 10 training epochs, and a loss function of CRF negative log-likelihood loss; an early stopping strategy is adopted during training, terminating training when the validation set F1-score does not improve for 3 consecutive epochs. The key parameters of the brand entity recognition model include: hidden layer dimension 768, number of attention heads 12, number of Transformer layers 12, vocabulary size 21128, dropout rate 0.1, maximum sequence length 512, number of label categories 5, CRF transition matrix dimension 5×5, learning rate 2e-5, batch size 32, and number of training epochs 10.
[0048] Step 5: Large Language Model-Assisted Disambiguation. For ambiguous candidate entities with confidence scores between 0.4 and 0.7, the large language model is invoked for contextual semantic disambiguation. If the large language model (LLM) determines the entity as a brand, the confidence score is weighted at 0.15; otherwise, the confidence score is multiplied by 0.5. The large language model employs a Transformer decoder-only architecture, consisting of multiple multi-head self-attention mechanisms, a feedforward neural network, and layer normalization. Each self-attention layer contains a multi-head attention mechanism, achieving autoregressive generation through causal masks; the feedforward network consists of two layers of linear transformations and a GELU activation function. The model input, after word segmentation, is mapped to word embedding vectors and superimposed with positional encodings. After layer-by-layer computation, the vectors are projected onto the vocabulary space through the output layer, generating output text token-by-token in an autoregressive manner.
[0049] Specifically, the training process of this large language model is as follows: First, unsupervised pre-training is performed on over 500GB of Chinese internet corpus, using an autoregressive language modeling objective, i.e., predicting the next token based on the preceding token. After pre-training, an instruction fine-tuning dataset containing a brand entity disambiguation task is constructed, and supervised fine-tuning (SFT) is used to teach the model to judge whether a candidate entity is a brand entity based on the context. Further, reinforcement learning based on human feedback (RLHF) is used to align the model output with preferences, improving the accuracy of disambiguation judgments. Training uses the AdamW optimizer with cosine annealing learning rate scheduling. The key parameters of this large language model include: 32 Transformer layers, 4096 hidden layer dimensions, 32 attention heads, 65024 vocabulary words, 4096 maximum context lengths, 0.0 Dropout rate, 3e-4 pre-training learning rate, 1e-5 fine-tuning learning rate, 128 batch size, over 500GB of training data in Chinese corpus, and approximately 100,000 instruction fine-tuning data words.
[0050] Step Six: Multi-dimensional Confidence Calculation. The system outputs the four-tuple information for each candidate entity: First, the entity name EntityName; Second, the text position TextPosition, represented as a left-closed, right-open interval from start to end; Third, the context snippet ContextSnippet, taking N=50 characters before and after; Fourth, the recognition confidence, calculated using the formula: Confidence=w1×S_dict+w2×S_ctx+w3×S_ner+w4×S_llm. Here, S_dict is the dictionary matching score, with 1 for exact matching, 0.8 for edit distance of 1, 0.6 for edit distance of 2, and 0.5 for pinyin matching; S_ctx is the context rule score, adding 0.15 for positive rules and subtracting 0.2 for negative rules, and cropping to the range of 0 to 1; S_ner is the NER model probability, taking 0.5 when no model is used; S_llm is the LLM disambiguation score, taking 1 for brand identification, 0.2 for non-brand identification, and 1 when not invoked. The default weights are w1=0.35, w2=0.25, w3=0.25, and w4=0.15. This confidence calculation method, through weighted fusion of multi-source heterogeneous features, differs from existing technologies that solely rely on dictionary matching or model output.
[0051] Step 7: Determining the Confidence Threshold and Handling Low-Confidence Entities. The preset threshold, Threshold, is determined based on the confidence value obtained when the F1-score is optimally adjusted using a standard test set. This standard test set contains over 500 labeled AI response samples, with an empirical default value of 0.65. Users can adjust it in 0.05 increments within the range of 0.40 to 0.85. Low-confidence entities are weighted using a confidence ratio reduction method. The adjusted weight = original weight × Confidence / Threshold. For example, an entity with a confidence level of 0.45 and a threshold of 0.65 will have its statistical weight reduced to 0.45 / 0.65, approximately 0.69.
[0052] Specifically, the optimal harmonic evaluation value F1-score refers to classifying candidate entities as brand entities or non-brand entities on the standard test set using different confidence values as classification thresholds, calculating the corresponding precision and recall respectively, and then calculating the harmonic evaluation value F1-score = 2 × Precision × Recall / (Precision + Recall). After traversing all candidate thresholds, the confidence value corresponding to the maximum F1-score is taken as the preset threshold. The specific process is as follows: Step 1: Prepare a standard test set containing over 500 manually annotated AI response samples. Each sample's brand entities have been independently annotated by at least two annotators with consistent results. Step 2: Perform brand entity recognition on each response in the test set, outputting the confidence score for each candidate entity. Step 3: Iterate from 0.40 to 0.85 in increments of 0.05. For each candidate threshold T, entities with a confidence score not lower than T are considered positive examples. Calculate the corresponding Precision and Recall, where Precision is the ratio of correctly identified brand entities to the total number of identified brand entities, and Recall is the ratio of correctly identified brand entities to the total number of actual brand entities in the test set. Step 4: Calculate the F1-score for each threshold. Select the threshold with the highest F1-score as the system default threshold, with an empirical value of 0.65. This threshold can be manually adjusted by the user based on business scenarios after system deployment. For candidate entities with confidence levels below a preset threshold, the system can mark them as pending confirmation and not directly include them in the core share; alternatively, the system can reduce their statistical weight based on user configuration.
[0053] Specifically, the preset threshold, Threshold, is determined based on the confidence value at which the F1-score is optimal on a standard test set containing over 500 labeled AI response samples. The empirical default value is 0.65, which users can adjust in 0.05 increments within the range of 0.40 to 0.85. Low-confidence entities are weighted proportionally based on their confidence level; the adjusted weight is calculated as: Original Weight × Confidence / Threshold. For example, an entity with a confidence level of 0.45 and a threshold of 0.65 has its statistical weight reduced to 0.45 / 0.65, approximately 0.69 times. This method differs from existing techniques that directly discard low-confidence entities or simply reduce weight proportionally. It achieves smooth weighting through the ratio of confidence level to threshold, preserving the information contribution of low-confidence entities while avoiding their excessive influence on share calculation.
[0054] In this optional embodiment, the brand entity configuration includes a standard brand code, a standard brand name, abbreviation, product name, series name, historical name, official website domain name, corporate entity, and exclusion words. Standard brand merging is performed on candidate brand entities to obtain valid brand mentions bound to the standard brand code. This includes: performing precise name matching, product name attribution matching, official website domain name association, and corporate entity association on the candidate brand entities sequentially to obtain the brand merging result; when the brand merging result points to a unique standard brand code, the corresponding candidate brand entity is bound to the unique standard brand code; when product names are duplicated across brands or the brand entity relationship is unclear, conflict resolution is performed based on the confirmed brand co-occurrence relationship in the context; when conflict resolution fails, the corresponding candidate brand entity is marked as pending confirmation; and candidate brand entities that are bound to the unique standard brand code, appear in the valid content area of the response text, and whose identification confidence meets the preset inclusion threshold are determined as valid brand mentions.
[0055] (iv) Standard Brand Consolidation Mechanism: Specifically, such as Figure 3 As shown, the standard brand merging in this invention is used to solve the problem of multiple expressions for the same brand. It merges the Chinese name, English name, abbreviation, historical name, product name, series name, official website name, and corporate name of the same brand into a single standard brand.
[0056] Step 1: Load the Brand Configuration Table. Load the brand entity configuration table from the system configuration database. This brand entity configuration table includes fields such as BrandID primary key, standard Chinese name, English name, a list of abbreviations in JSON format, a list of product names including codes and names, series name, historical name, official website domain name, enterprise entity, and excluded words. Each BrandID represents an indivisible brand statistical unit.
[0057] Step 2: Hierarchical Merging and Matching. Hierarchical merging and matching employs a four-level priority pipeline. Priority 1 is exact name matching: candidate entity names are merged directly when they are exactly equal to the brand's standard Chinese name, English name, or abbreviation. Priority 2 is product name attribution matching: candidate entity names are merged to their respective brands when they are exactly equal to any product in the product name list. Priority 3 is website domain association: candidate entities are merged to their corresponding brands when their context includes a configured website domain. Priority 4 is enterprise entity association: candidate entities are merged when their enterprise matches the brand's configured entity.
[0058] Step 3: Conflict Resolution. When product names are duplicated across brands, disambiguation is performed by confirming the co-occurrence relationship of the brands through context; if resolution fails, it is marked as pending confirmation and is not included in the core share.
[0059] Step Four: Definition of Three Characteristics of Brand Mentions. A valid brand mention, as defined in this invention, must simultaneously meet three technical characteristics: First, a valid merge mapping is established between the text fragment and a standard brand code in the system's brand configuration table; second, it appears in a valid content area of the AI response, which does not include URL strings, AI model signatures, disclaimers, or purely formatting tags; and third, the final confidence level is not lower than a preset inclusion threshold. Only when all three technical characteristics are met simultaneously is it counted as a valid brand mention and included in subsequent share calculations.
[0060] For text fragments containing common nouns, place names, personal names, industry concepts, or excluded words, the system will not count them as brand mentions. In cases where product names are duplicated across brands or the brand relationship is unclear, the system may mark them as pending confirmation, reduce their weight, or merge or split them according to the user-configured relationship.
[0061] Specifically, regarding the aforementioned brand mentions, the following supplementary explanation is provided: A valid brand mention as defined in this invention must simultaneously meet three technical characteristics: First, a valid merge mapping is established between the text fragment and a standard brand code (BrandID) in the system's brand configuration table. This means that after two steps of brand entity identification and standard brand merging, it is successfully bound to a unique brand statistical unit. Second, it appears in the valid content area of the AI response, excluding URL strings, AI model signatures, disclaimer paragraphs, and purely formatted tags. Purely formatted tags include Markdown headings and separators. Third, the final confidence level is not lower than a preset inclusion threshold, which is 0.65 in this embodiment. Only when all three technical characteristics are met is it counted as a valid brand mention and included in subsequent share calculations. For cases where product names are duplicated across brands or the brand entity relationship is unclear, the system can mark it as pending confirmation and reduce its weight, or merge or split it according to the user-configured entity relationship. This entity relationship includes priority based on ownership or priority based on splitting, which differs from the simple binary processing method of full or discard in existing technologies.
[0062] Specifically, regarding the determination of the aforementioned preset inclusion threshold of 0.65, the following supplementary explanation is provided: This default value of 0.65 was determined based on the following experimental verification process. Threshold search was performed on a validation set containing 1200 labeled AI response samples. This validation set covers 8 mainstream AI platforms and 12 industry categories. The experimental results show that: at a threshold of 0.60, F1=0.82, Precision=0.78, Recall=0.87; at a threshold of 0.65, F1=0.85, Precision=0.84, Recall=0.86; and at a threshold of 0.70, F1=0.83, Precision=0.88, Recall=0.79. The threshold of 0.65 achieves the optimal balance between precision and recall, and is therefore used as the system default value. Users can adjust it themselves within the range of 0.40 to 0.85 in 0.05 increments. For scenarios requiring high integrity in brand monitoring, the threshold can be appropriately lowered, for example, adjusted to 0.55; for scenarios requiring high accuracy in share calculation, the threshold can be appropriately increased, for example, adjusted to 0.75.
[0063] Specifically, after the standard brands are merged, each valid brand mention is bound to a unified standard brand code, so as to facilitate unified counting, share calculation and multi-brand comparison in the future.
[0064] In this optional embodiment, the classification record includes a main category code, a list of secondary category codes, confidence scores for each category, user decision-making influence scores, and question weights. Determining question weights based on the classification record includes: generating category term variations based on the user-configured product or service name; expanding purchase questions based on the category term variations and generating reputation verification questions; reputation verification questions include brand names; extracting keywords and semantically classifying the question texts to which valid brand mentions belong, determining the main category code, a list of secondary category codes, and confidence scores for each category; calculating the user decision-making influence score based on question coverage, purchase intent strength, and position sensitivity; and determining the question weight based on the proportion of the user decision-making influence score in all question categories.
[0065] (v) Problem stratification and weight allocation mechanism: The system can categorize issues into categories such as category awareness, purchase recommendations, competitor comparison, reputation verification, risk avoidance, scenario requirements, price budget, and functionality / performance. The specific method for categorizing issues includes the following steps: Step 1: Intelligent Mining of Category Terms. The system automatically mines category term variations from three dimensions: professional terminology, common names, and colloquial terms. Taking a new media data analysis platform as an example, the professional terminology dimension produces new media data monitoring platforms, content data analysis SaaS tools, etc.; the common names dimension produces new media operation tools, self-media analysis software, etc.; and the colloquial terms dimension produces self-media data query tools, blogger data analysis tools, etc. After comprehensively judging user search frequency, user base breadth, and general recognition, 15 category term variations are selected.
[0066] Step Two: Dual-Channel Expansion of Purchase Questions. The left channel aggregates high-frequency question words from Baidu search engine, social media platforms, e-commerce platforms, e-commerce shopping guide platforms, and dropdown keywords in real time to obtain real user question data; the right channel adopts a dual-dimensional decision coverage matrix, with four decision-maker identities as the horizontal axis and six core decision points as the vertical axis, performing 24 cross combinations, and AI intelligently simulates and generates long-tail question words. The four decision-maker identities include personal use, gifting, purchasing on behalf of others, and procurement; the six core decision points include colloquial expression, scenario adaptation, functionality, budget, evaluation, and risk. For example, choosing a good content data analysis tool for an individual doing self-media corresponds to the combination of personal use and functionality, while choosing a cost-effective new media monitoring system for a team corresponds to the combination of procurement and budget.
[0067] Step 3: Automatic Generation of Reputation Questions. The system automatically generates four types of verification questions containing the brand name, including data query accuracy, operational tool usability, industry ranking reference, and business cooperation suitability, to monitor the accuracy and sentiment of the information in the AI's responses.
[0068] Step 4: Question Category Classification. Keywords are extracted from the question text using the TextRank algorithm and then compared with predefined feature word libraries for each category using Jaccard similarity calculation. For example, the word library for shopping recommendations includes "recommendation," "ranking," "which is good," "cost-effectiveness," "worth buying," and "recommendation"; the word library for competitor comparison includes "comparison," "difference," "vs," "review," and "comparative review"; the word library for reputation verification includes "how is it," "good or bad," "evaluation," and "reputation"; and the word library for price and budget includes "price," "how much," "cost-effective," and "cost-effectiveness." Boundary questions are classified using Sentence-BERT semantic classification; when matching multiple categories, the main category with the highest similarity is used for statistics, and switching between secondary categories is supported. Simultaneously, the system automatically filters obviously irrelevant questions, such as "Can data from content analysis be fake?", to ensure the purity of the question set.
[0069] Step 5: Importance Weight Configuration. The user decision-making influence score method is adopted: InfluenceScore(Category) = Reach × Intent × Position. Where, Reach is the coverage rate, taken as the normalized value of the estimated monthly search volume for this type of question, weighted at 0.70 for real data in the left channel and 0.30 for simulated data in the right channel; Intent is the intensity of purchase intent, scored from 0 to 1.00 based on the density and intensity of purchase intent keywords, adding 0.15 for purchase selection keywords such as "recommendation" and "which is better," and adding 0.10 for comparison keywords such as "comparison" and "review"; Position is the sensitivity to ranking, quantified based on the difference in impact of the brand's ranking in the recommendation list on user click-through rate, with the first position assigned 1.00, the top three 0.75, and subsequent positions 0.45. The final weight is Weight = InfluenceScore / ΣInfluenceScore. This method differs from existing technologies that use subjective assignment or equal weighting. The result of the hierarchical question processing is a classification record generated for each question, containing fields such as QuestionID, question text, main category code, a list of secondary category codes in JSON format, confidence score for each category, InfluenceScore, normalized weight, and classification method identifier. The classification method identifier includes keywords, semantics, or a combination thereof. This result serves as the input parameter for the weighted calculation of the overall AI brand market share, ensuring that brand share reflects not only the frequency of occurrence but also the difference in decision-making value between question categories. This mechanism makes brands visible at the category awareness level and allows brands to be chosen at the decision-making level.
[0070] Regarding the aforementioned methods for generating category term variations, this invention employs a three-dimensional automatic category term mining method, with the following specific steps: Step 1: Input Processing: Use the user-configured product name or service name as the seed term, such as "new media data analysis platform." Step 2: Professional Terminology Mining: Expand the seed term with industry terminology using a large language model, constraining the output to be formal technical names within the same functional domain, such as "new media data monitoring platform," "content data analysis SaaS tool," and "social media sentiment analysis system," generating 5 to 8 candidate terms each time. Step 3: Common Terminology Mining: Based on search engine dropdown keywords and related search interfaces, access Baidu search engine, social media platforms, e-commerce platforms, e-commerce shopping guide platforms, and dropdown keywords to capture high-frequency user search terms semantically related to the seed term, such as "new media operation tool" and "self-media analysis software," selecting the Top-10 terms in descending order of search frequency within 30 days. Step 4: Colloquialism Dimension Mining: Utilize a large language model to simulate colloquial expressions used by ordinary users, generating informal but frequently used terms, such as those used in self-media data query tools and blogger data analysis tools. The output should be colloquial, concise, and reflective of user feedback. Step 5: Comprehensive Screening: All candidate words generated from the three dimensions are comprehensively scored, typically producing 20 to 30 candidate words. The comprehensive score is determined based on three factors: user search frequency, user reach, and general recognition. User search frequency is the normalized cumulative search volume of the word across various platforms over 30 days, with a weight of 0.50; user reach is the ratio of the number of platforms covered by the word to the total number of platforms, with a weight of 0.30; and general recognition is the probability that the large language model determines whether ordinary users can understand the word, with a weight of 0.20. Step 6: Sort by comprehensive score in descending order, and select the Top-15 as the final category word variant set output.
[0071] Specifically, the two-dimensional decision coverage matrix is explained below.
[0072] I. Specific Definitions of the Four Decision-Maker Identityes: Identity 1 is the self-use decision-maker, referring to the end-user who purchases or uses products or services for themselves. Their questioning characteristics include a focus on personal experience, convenience, and cost-effectiveness, such as choosing the best tools for personal social media creation. Identity 2 is the gift-giving decision-maker, referring to the purchaser who selects products for others. Their questioning characteristics include a focus on brand awareness, packaging, and the recipient's preferences, such as choosing a coffee brand that is prestigious to give as a gift to a friend. Identity 3 is the proxy purchase decision-maker, referring to the executor who makes purchases on behalf of others. Their questioning characteristics include a focus on specific specifications and models, and clear indications, such as choosing the best-selling coffee brand to buy for a colleague. Identity 4 is the procurement decision-maker, referring to the person in charge of bulk purchases for a company or team. Their questioning characteristics include a focus on bulk discounts, supply stability, and compliance, such as choosing the most cost-effective new media monitoring system for a team.
[0073] II. Specific Definitions of the Six Core Decision Points: Decision Point 1 is conversational expression, where users describe their needs in everyday language without mentioning specific brands or parameters, such as "delicious coffee" or "useful tools"; Decision Point 2 is scenario adaptation, where users describe specific usage scenarios and focus on the product's compatibility with those scenarios, such as "drinking it on the commute" or "using it in the office"; Decision Point 3 is functionality, where users focus on the product's core functions and technical specifications, such as support for multi-platform data and the ability to generate analysis reports; Decision Point 4 is budget, where users focus on price range and cost-effectiveness, such as "coffee under 500 yuan per month" or "coffee under 20 yuan per month"; Decision Point 5 is evaluation, where users focus on feedback and word-of-mouth from other users, such as what users say and whether the product has a good reputation; Decision Point 6 is risk, where users focus on potential negative factors and information to avoid pitfalls, such as what pitfalls exist and what to do if the service is inadequate.
[0074] III. Cross-combination Implementation Method: The system constructs a 4×6=24-cell cross-matrix. For each cell, using the user-configured category terms as input, the large language model generates question terms according to the following Prompt template: From the perspective of a decision-maker, describe the decision-making points of the category terms and generate three natural, conversational questions. Each cell generates three question terms, resulting in 72 candidate question terms across 24 cells. After deduplication and semantic similarity filtering, the final question term set is output. Candidate question terms with a cosine similarity greater than 0.85 are considered duplicates; only the one with the higher search frequency is retained.
[0075] Specifically, the present invention uses a multi-dimensional automatic classification method based on the features of the question text to determine the identity of the decision-maker. The specific steps are as follows.
[0076] Step 1: Feature Terminology Building: Establish a dedicated feature terminology library for each decision-maker identity. The feature terminology library for self-use decision-makers includes first-person expression of needs such as "for myself," "personally," "I want," and "suitable for me." The feature terminology library for gift-giving decision-makers includes gift-giving, "gift," "prestigious," "good-looking packaging," "gift for friends," and "gift for elders." The feature terminology library for purchasing on behalf of others includes delegation-execution terms such as "help," "on behalf of," "colleagues asked me to buy," "someone asked me to buy," and "which one has the highest sales." The feature terminology library for procurement decision-makers includes organizational procurement terms such as "team," "company," "bulk," "procurement," "enterprise version," "commercial," and "centralized procurement."
[0077] Step 2, Feature word hit scoring: After segmenting the input question text, the number of hits in the feature word library for each identity and the sum of the TF-IDF weights of the hit words are counted to obtain the FeatureScore(Identity) for each identity.
[0078] Step 3: Sentence Template Matching: Predefine typical sentence templates for each identity. For self-use: I + verb + what + category word + good; for gift-giving: give + recipient + what + category word; for purchasing on behalf: help + recipient + verb + category word; for procurement: company or team + verb + category word + which company. Perform dependency parsing on the query text and calculate the structural similarity score (StructScore(Identity)) with each template.
[0079] Step 4, Comprehensive Judgment: The final identity determination score is IdentityScore = 0.6 × FeatureScore + 0.4 × StructScore. The identity with the highest score is taken as the determination result. When the difference between the highest and second-highest scores is less than 0.15, it is marked as an ambiguous identity, and the system processes it as a default identity, and the confidence level is marked in the classification record. This method uses a dual-channel fusion of feature word matching and sentence structure analysis to ensure the accuracy and feasibility of the decision-maker's identity determination.
[0080] Specifically, the present invention employs a two-layer cascaded method of feature lexicon and semantic classifier to automatically identify the core decision point category to which the question text belongs, which includes the following steps.
[0081] Step 1: Constructing a feature vocabulary library for six decision points: The colloquial expression feature vocabulary library includes words like "useful," "delicious," "reliable," "which is good," "what's good," "recommend," and "do you have any"; the scenario-adaptation feature vocabulary library includes words like "commuting," "office," "business trip," "outdoor," "home," "party," "morning," "evening," and "going to work"; the function feature vocabulary library includes words like "support," "function," "can / cannot," "can," "how to use," "multi-platform," "analysis," "monitoring," "report," and "export"; the budget feature vocabulary library includes words like "price," "how much," "expensive," "cost-effective," "value for money," "cheap," "budget," "monthly fee," "annual fee," and "free"; the evaluation feature vocabulary library includes words like "how is it," "good / bad," "reputation," "evaluation," "experience," "users," "feedback," "reliable," and "worth it"; and the risk feature vocabulary library includes words like "pitfall," "avoid pitfall," "disadvantages," "bad," "bad reviews," "problems," "risks," "after-sales service," "refund," and "complaint."
[0082] Step 2, Feature Word Fast Matching Layer: Segment the input question text, count the number of hits in the feature word library for each decision point and the TF-IDF weighted score, and output a six-dimensional feature vector FeatureVec=[F_oral,F_scene,F_func,F_budget,F_review,F_risk].
[0083] Step 3, Semantic Classifier Fine-tuning Layer: For boundary problems where feature word matching results are unclear, the query text is input into a six-class semantic model fine-tuned based on Sentence-BERT, and the output is the semantic probability distribution of each decision point: SemanticProb=[P_oral,P_scene,P_func,P_budget,P_review,P_risk]. Here, a boundary problem refers to a problem where the difference between the highest and second-highest feature scores is less than 0.2.
[0084] Step 4, Fusion Decision: The final decision point score is calculated as DecisionScore = 0.55 × Normalize(FeatureVec) + 0.45 × SemanticProb. When the feature word match is clear, only the first layer result is used, skipping the semantic classifier to improve efficiency. The class with the highest score is used as the primary decision point, and the class with the second highest score (greater than 0.25) is used as the secondary decision point.
[0085] Step 5, Quality Verification: When the scores for all categories are less than 0.30, the issue is marked as having an unclear decision point and enters the manual review queue or is categorized according to the default category's colloquial expression. This two-layer cascaded method differs from existing technologies that rely solely on keyword matching or deep learning models. Through a hierarchical architecture of rapid feature word filtering and semantic ranking, it controls computational overhead while ensuring classification accuracy.
[0086] Different question categories have varying values for brand operations, therefore the system supports assigning importance weights to different question categories. For example, purchase recommendation questions can be assigned higher weights, word-of-mouth verification questions can be used to focus on analyzing sentiment trends, and category awareness questions can be used to measure the brand's presence in the early decision-making stage.
[0087] The results of the problem stratification are used in the subsequent comprehensive AI brand market share calculation, so that the brand share not only reflects the number of occurrences, but also reflects the difference in the value of the problem.
[0088] Specifically, the result of the hierarchical processing of questions in this invention generates a classification record, QuestionClassificationRecord, for each question, containing the following fields: QuestionID (unique identifier of the question), the original text of the question, the main category code (CategoryCode), such as PURCHASE_REC, COMPETITOR_COMPARE, REPUTATION_VERIFY, and PRICE_BUDGET; the list of secondary category codes is in JSON array format, such as REPUTATION_VERIFY and SCENE_REQUIRE; and the confidence level for each category. Scores, ranging from 0 to 1.00; InfluenceScore, calculated using the user decision-making influence scoring method, where InfluenceScore is the product of Reach, Intent, and Position; Normalized Weight, Weight = InfluenceScore / ΣInfluenceScore, ensuring the sum of weights for all question categories is 1.00; and a Classification Method Flag, with values of KEYWORD, SEMANTIC, or HYBRID, indicating whether the question is classified using keyword matching, semantic classification, or a hybrid method. This classification record serves as input parameters for the weighted calculation of comprehensive AI brand market share, ensuring that brand share reflects not only frequency of occurrence but also the difference in decision-making value between question categories. For example, the Weight for competitor comparison questions is typically higher than that for category awareness questions because the former is closer to the user's final purchase decision stage.
[0089] In this optional embodiment, determining the appearance position, recommendation ranking, and position weight includes: parsing the answer text to obtain the recommendation list area, the main text paragraph area, the supplementary explanation area, and the cited source area; when a valid brand mention is located in the recommendation list area, determining the recommendation ranking according to the list arrangement order; when a valid brand mention is not located in the recommendation list area, determining the appearance position according to the positional relationship between the paragraph containing the valid brand mention and the beginning, middle, or end of the answer text; determining the position weight based on the user attention decay relationship corresponding to the appearance position; wherein, the position weight is determined according to the decreasing display priority in the answer text.
[0090] Specifically, the user attention decay relationship refers to the quantitative correspondence that the degree of user attention to content displayed in different positions in the AI answer decreases progressively as the position moves forward. This decay relationship is derived from the click-through rate distribution statistics of user behavior data of AI answers on multiple platforms. Specifically, the click-through rate and reading completion rate are highest for high-priority display areas such as the first position in the recommended list and the first paragraph of the main text. As the display position moves forward in sequence to the later positions in the recommended list, the middle of the main text, the last paragraph of the main text, the supplementary explanation area, and the citation source area, the user's click-through rate and attention decrease progressively. The system quantifies this decay relationship into a position weight coefficient, with the user click-through rate of the first position in the recommended list as the baseline value of 1.00. The weight of other positions is determined by the ratio of their actual click-through rate to the baseline click-through rate, forming a decreasing weight sequence from 1.00 to 0.10.
[0091] In this optional embodiment, the display priority is as follows: the first item in the recommended list is higher than the second and third items in the recommended list; the second and third items in the recommended list are higher than the first paragraph of the main text; the first paragraph of the main text is higher than the fourth item and thereafter in the recommended list; the fourth item and thereafter in the recommended list are higher than the middle paragraphs of the main text; the middle paragraphs of the main text are higher than the last paragraph of the main text; the last paragraph of the main text is higher than the supplementary explanation area; and the supplementary explanation area is higher than the cited source area.
[0092] In this optional embodiment, the AI brand market share includes general mention share, answer coverage share, question coverage share, position-weighted share, sentiment-weighted share, and comprehensive AI brand share. The determination of the AI brand market share includes: calculating the general mention share based on the number of valid brand mentions corresponding to each standard brand code; analyzing the answer coverage share based on the number of answers that mention the corresponding standard brand code at least once; determining the question coverage share by the number of questions that have appeared under the corresponding standard brand code; solving for the position-weighted share using the position weights of each valid brand mention under the corresponding standard brand code; generating the sentiment-weighted share based on the sentiment tendency of each valid brand mention under the corresponding standard brand code; and weighting and fusing the general mention share, answer coverage share, question coverage share, position-weighted share, and sentiment-weighted share to obtain the comprehensive AI brand share.
[0093] In this optional embodiment, the multi-brand comparative analysis results include an overview of AI brand market share, a multi-brand ranking, question details, platform comparison, trend analysis, sentiment comparison, and shortcomings. The AI brand market share overview displays the product's overall AI brand share, general mention share, position-weighted share, and question coverage share compared to competitors. The multi-brand ranking is sorted by overall AI brand share, mention count, average ranking, and positive mention ratio. Question details display the mention status, ranking, share, and answer snippets for each question for both the product and competitors. Platform comparison displays the share and ranking differences between the product and competitors on different AI platforms. Trend analysis displays the share changes of the product and competitors over a time period. Sentiment comparison displays the positive, negative, and neutral mention ratios for both the product and competitors. Shortcomings identify issues that the product does not address, has a low ranking, a high negative ratio, or where competitors have a clear advantage.
[0094] (vi) Location and Recommendation Ranking Identification Mechanism: The system identifies the position and recommendation ranking of the standard brand in each answer. Positions include first place in the recommendation list, top three places in the recommendation list, first paragraph of the main text, middle paragraph of the main text, last paragraph of the main text, question and answer paragraph, quotation paragraph, and supplementary explanation paragraph.
[0095] Specifically, the location identification method includes the following steps.
[0096] Step 1: Answer Structure Analysis. The AI's answer text is analyzed using regular expression matching to identify the recommendation list area, main body paragraph area, supplementary explanation area, and source citation area. The recommendation list area can be identified using numerical codes and keywords indicating recommendation intent such as "recommended," "preferred," "worthwhile," and "TOP." The supplementary explanation area can be identified using guiding words such as "note" and "supplementary." The source citation area can be identified using URLs and reference markers.
[0097] Step 2: Calculate the recommended ranking. In the recommended list area, assign recommended rankings such as Rank=1, Rank=2, Rank=3, etc., according to the order of arrangement; for non-list formats, infer the implicit rankings using sequence relation words.
[0098] Step 3: Categorize Appearance Locations. Appearance locations are categorized into 8 types: first position in the recommended list, top three positions in the recommended list, subsequent positions in the recommended list, first paragraph of the main text, middle paragraph of the main text, last paragraph of the main text, supplementary explanation paragraph, and cited explanation paragraph. Location weight configuration uses a user click-through rate decay model, Wpos = BaseClickRate / BaseClickRate_First. Specific weights are: first position in the recommended list = 1.00, top three positions = 0.75, subsequent positions in the recommended list = 0.45, first paragraph of the main text = 0.60, middle paragraph of the main text = 0.35, last paragraph of the main text = 0.25, supplementary explanation paragraph = 0.15, and cited explanation paragraph = 0.10. Default weights can be customized by the user.
[0099] The system can assign different weights to different positions. For example, if a brand appears at the top or top three of the recommendation list, it indicates that it has a strong influence on users' decisions; if a brand only appears in the supplementary descriptions or at the end, its influence weight is lower.
[0100] Specifically, the rules for configuring position weights are as follows: This invention uses a user click-through rate decay model to assign weights to different positions, calculated as Wpos = BaseClickRate / BaseClickRate_First. The first position in the recommendation list has a weight of 1.00, serving as the baseline and corresponding to the highest user attention and click-through rate; the first three positions have a weight of 0.75, as users typically view these positions, but attention has already decayed; subsequent positions have a weight of 0.45, as users need to scroll or turn pages to see them, resulting in a significant drop in click-through rate; the first paragraph of the main text has a weight of 0.60, as the first paragraph of an AI response is usually a summary and has a high readability; the middle section of the main text has a weight of 0.35, as user reading depth decreases; the last paragraph of the main text has a weight of 0.25, as some users will not read to the end; the supplementary explanation section has a weight of 0.15, usually AI-added notes with low user attention; and the citation section has a weight of 0.10, simply for references or source citations. The final position-weighted contribution for each brand mention is MentionCount × Wpos. The default weight values mentioned above are derived from statistical analysis of user behavior data from multiple platforms' AI-generated answers. This data includes over 10,000 user click heatmaps. Users can also customize the weight value for any position type within the range of 0 to 1.00 according to their own business scenarios. Through position weights, this invention can distinguish between being mentioned and being highlighted.
[0101] (vii) Mechanism for calculating market share of AI brands: It should be noted that, as Figure 4As shown, the market share of AI brands in this invention includes general mention share, answer coverage share, question coverage share, location-weighted share, sentiment-weighted share, and comprehensive AI brand share.
[0102] The common mention share, Share_count(B), can be expressed as: Share_count(B) = Mention(B) / ΣMention(i). Here, Mention(B) represents the aggregated mention count of brand B across all valid AI responses, and ΣMention(i) represents the sum of aggregated mention counts of all followed brands.
[0103] The answer coverage share, Share_answer(B), can be expressed as: Share_answer(B) = AnswerHit(B) / ΣAnswerHit(i). Here, AnswerHit(B) represents the number of valid answers that mention brand B at least once.
[0104] The question coverage share, Share_question(B), can be expressed as: Share_question(B) = QuestionHit(B) / TotalQuestion. Where QuestionHit(B) represents the number of valid questions that brand B has had, and TotalQuestion represents the total number of questions within the analysis scope.
[0105] The position-weighted share Share_pos(B) can be expressed as: Share_pos(B) = ΣWpos(B,r) / ΣΣWpos(i,r). Where Wpos(B,r) represents the position weight of brand B in response r.
[0106] The sentiment weighting value can be expressed as Wsent = 1 + α × Positive - β × Negative. Here, Positive represents a positive sentiment marker, Negative represents a negative sentiment marker, and α and β are configurable parameters.
[0107] The overall AI brand share, AIShare(B), can be expressed as: AIShare(B) = a × Share_count + b × Share_answer + c × Share_question + d × Share_pos + e × Share_sent. Where a, b, c, d, and e are configurable weights.
[0108] (viii) Anomaly Correction Mechanism: The system performs anomaly correction for low-confidence brand candidate entities, duplicate responses, brand ambiguity, cross-brand product name duplication, and multiple brands sharing the same entity.
[0109] Specifically, anomaly correction methods include automatic filtering, reducing weight, marking as pending confirmation, merging or splitting the subject. For highly similar duplicate answers to the same task, the same question, the same platform, and the same time window, the system can retain one or reduce the weight of the duplicate according to deduplication rules.
[0110] Specifically, when a user adjusts the competitor collection or brand merging relationship, the system can recalculate historical reports or adopt a new standard from the adjustment point in time and indicate the change in standard in the report.
[0111] (ix) Multi-brand comparative analysis and report output: The system outputs a multi-brand comparative analysis, including an overview of AI brand market share, multi-brand rankings, problem details, platform comparison, trend analysis, sentiment comparison, shortcomings, and indicator explanations.
[0112] Specifically, the AI Brand Market Share Overview displays the overall AI market share, general mention share, position-weighted share, and question coverage share of this product and its competitors; the Multi-Brand Ranking sorts products by indicators such as overall share, number of mentions, average ranking, and positive mention ratio; and the Question Details display the mention status, ranking, share, and answer snippets of this product and its competitors for each question.
[0113] Specifically, the platform comparison is used to display the differences in brand share and ranking across different AI platforms; the trend analysis is used to display changes in brand share within the selected period; the sentiment comparison is used to display the positive, negative, and neutral mention ratios of this product and its competitors; and the weakness issues are used to identify problems that this product does not have, that are ranked low, that have a high negative ratio, or that have obvious advantages over its competitors.
[0114] Figure 5 and Figure 6 An embodiment of an AI brand market share analysis system of the present invention is shown.
[0115] In this alternative embodiment, such as Figure 5 As shown, this AI-powered brand market share analysis system includes a configuration and access layer, an identification and merging layer, and an analysis and output layer. Figure 6As shown, the AI brand market share analysis system specifically includes: a brand entity configuration module 201, used to generate analysis tasks based on user configuration, including the product itself, competitors, brand entity configuration, and analysis scope configuration; a question scope configuration module 202, used to read AI answer records according to the analysis task and determine the set of answers to be analyzed; a brand entity recognition module 203, used to perform brand entity recognition on the answer text in the set of answers to be analyzed, obtaining brand candidate entities; the brand candidate entities include entity name, text position, context fragment, and recognition confidence; and a standard brand merging module 204, used to perform brand candidate entity merging based on the brand entity configuration. The system performs standard brand merging to obtain valid brand mentions bound to standard brand codes; the question stratification module 205 generates classification records based on the question text to which valid brand mentions belong, and determines the question weight based on the classification records; it determines the occurrence position, recommendation position, and position weight by the text position of valid brand mentions in the answer text, and identifies the sentiment tendency corresponding to valid brand mentions; the market share calculation module 206 calculates the AI brand market share of this product and competitors based on valid brand mentions, question weight, position weight, and sentiment tendency; and the comparative analysis module 207 generates multi-brand comparative analysis results based on the AI brand market share.
[0116] To facilitate understanding of the above technical solution of the present invention, the following is a specific explanation using a brand analysis scenario as an example: Example 1: AI Market Share Comparison Analysis of Coffee Brands. A user creates a coffee brand subscription program, designating this product as coffee brand A, and configuring its English name, abbreviation, official website domain, and main product names. The user also configures coffee brands B, C, and D as competitors. The system reads the AI's response records from the past 30 days for this program, covering topics such as affordable coffee recommendations, commuter coffee choices, coffee delivery brands, coffee brand reputation, and avoiding common coffee product pitfalls.
[0117] The system identifies the brand entity in each answer and categorizes the Chinese name, English name, and star product name of coffee brand A under the "Coffee A" brand. For coffee brands B and C appearing in the answers, the system categorizes them by competitor entity. If a product name may refer to two brands simultaneously, the system marks that record as pending confirmation and excludes it from the core share.
[0118] After calculation, the system found that coffee brand A had a 28% share of general mentions, a 22% share of position-weighted mentions, and a 60% share of problem coverage; while coffee brand B had a 31% share of general mentions, a 35% share of position-weighted mentions, and a 64% share of problem coverage. Based on this, the system determined that although coffee brand A appeared almost as frequently as coffee brand B, it lagged behind in the recommendation ranking, and listed the recommendation of affordable coffees and the selection of coffees for commuting as key optimization issues.
[0119] Example 2: Multi-Platform Market Share Difference Analysis. Within the same brand project, the system calculates results based on AI platform segmentation. The report shows that Coffee Brand A has a higher overall market share on Platform A, but its share on Platform B is significantly lower than its competitors. Further examination of the issue details reveals that Platform B more frequently recommends competitors when answering price budget-related questions, and this product often appears in the supplementary description section rather than at the top of the recommended list. Based on this, the system outputs content optimization suggestions and an issue list for Platform B.
[0120] Example 3: Sentiment Share Analysis. After performing sentiment recognition on brand-related segments, the system found that Coffee Brand A ranked second in overall share, but its negative mention ratio was higher than other competitors. The report further shows that negative issues mainly focus on two dimensions: taste consistency and store service. Users can click on negative entries to view corresponding AI response segments and the source of the questions.
[0121] As can be seen from the above embodiments, the present invention can perform brand entity recognition, standard brand merging, question stratification, location recognition, emotion recognition, and market share calculation for the product and its competitors in AI answering scenarios, and generate multi-brand comparative analysis results based on the calculation results, thereby helping users understand the differences in brand performance between the product and its competitors under different questions, different AI platforms, different recommendation positions, and different emotional tendencies.
[0122] In addition, the present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0123] In one embodiment, the computer device may be a server, and its internal structure diagram may be as follows: Figure 7 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores static and dynamic information data. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the above method embodiments.
[0124] Those skilled in the art will understand that Figure 7The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the computer device to which the present invention is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0125] In addition, the present invention also provides a storage medium, which is a computer-readable storage medium storing a computer program, and the computer program, when executed by a processor, implements the steps in the above method embodiments.
[0126] In addition, the present invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0127] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0128] This invention is not limited to the structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this invention is limited only by the appended claims.
Claims
1. An AI-based brand market share analysis method, characterized in that, include: Based on user configuration, generate analysis tasks that include configurations for this product, competitors, brand entities, and analysis scope. Based on the analysis task, read the AI's response records and determine the set of responses to be analyzed; Brand entity recognition is performed on the answer text in the set of answers to be analyzed to obtain brand candidate entities; the brand candidate entities include entity name, text position, context fragment and recognition confidence; based on the brand subject configuration, the brand candidate entities are merged into standard brands to obtain effective brand mentions bound to standard brand codes; Based on the question text associated with the effective brand mentions, a category record is generated, and the question weight is determined based on the category record; By determining the text position of the effective brand mentions in the answer text, the occurrence position, recommendation position and position weight are determined, and the emotional tendency corresponding to the effective brand mentions is identified. Based on the effective brand mentions, the question weights, the position weights, and the sentiment biases, calculate the AI brand market share of this product and its competitors; based on the AI brand market share, generate multi-brand comparative analysis results.
2. The AI brand market share analysis method according to claim 1, characterized in that, The analysis task is generated based on the set of brands to focus on, the configuration of brand subjects, and the configuration of the analysis scope. The set of brands to watch includes the user-configured products and competing products; The brand entity configuration includes the brand alias, product name, official website domain name, corporate entity, and excluded keywords corresponding to the brand set; The analysis scope configuration includes the user-configured project scope, problem scope, AI platform scope, time period, and industry category; The brand entity configuration is used for standard brand merging; the question scope, the AI platform scope, and the time period are used to read AI answer records; and the industry category is used to match sample questions and answers from a public AI question and answer sample library.
3. The AI brand market share analysis method according to claim 2, characterized in that, The determination of the set of answers to be analyzed is performed through a three-level data acquisition mechanism, specifically including: Based on the analysis task, the first set of candidate answers is read from the AI answer records and subscription tracking results within the system authorized by the user. When the first candidate answer set is unavailable, read the second candidate answer set from the user-imported question and answer file; When both the first candidate answer set and the second candidate answer set are unavailable, a third candidate answer set is obtained by matching sample questions and answers from a public AI question and answer sample library based on the product, the competitor's product, the industry category, and the question scope. Based on the currently available data sources, determine the set of answers to be analyzed from the first set of candidate answers, the second set of candidate answers, or the third set of candidate answers, and configure the data source identifier and quality level for the set of answers to be analyzed.
4. The AI brand market share analysis method according to claim 3, characterized in that, The process of matching sample questions and answers from a publicly available AI question-and-answer sample library to obtain a third set of candidate answers includes: Based on the semantic similarity between the scope of the questions and the sample question texts, a question relevance score is obtained; A brand hit score is generated based on whether the sample response text matches the product or the competitor's product. The industry matching score is determined based on the degree of matching between the sample industry labels and the industry categories corresponding to the analysis task. A timeliness score is generated based on the proximity of the sample response time to the time period. The sample relevance score is obtained by weighting and fusing the score based on the question relevance, brand accuracy, industry matching, and timeliness. The sample questions and answers in the publicly available AI question and answer sample library are sorted according to their relevance, and a preset number of sample questions and answers at the top of the sort are selected as the third candidate answer set.
5. The AI brand market share analysis method according to claim 1, characterized in that, Brand entity recognition is performed on the answer text in the set of answers to be analyzed to obtain brand candidate entities, including: The response text is preprocessed to obtain a sentence sequence; The sentence sequence is matched using a multi-source dictionary based on a brand recognition dictionary system, and context rule matching is performed based on context fragments to obtain dictionary matching scores and context rule scores. The sentence sequence is sequence-labeled using a brand entity recognition model to obtain the entity recognition model probability; and initial candidate entities within a preset confidence interval are disambiguated to obtain a disambiguation score. The recognition confidence is determined based on the dictionary matching score, context rule score, entity recognition model probability, and disambiguation score, and the brand candidate entity is output.
6. The AI brand market share analysis method according to claim 5, characterized in that, The brand identification dictionary system includes a precise brand dictionary, a fuzzy alias dictionary, and an exclusion word dictionary; The precise brand dictionary is constructed based on the brand's standard name, English name, abbreviation, and product name, and is used for precise matching of the sentence sequence; The fuzzy alias dictionary is built based on common misspelled names, colloquial terms, and abbreviation variations, and is used to perform fuzzy matching on text fragments that do not match precisely. The exclusion word dictionary is built based on place names, common nouns, industry terms and user-defined exclusion words, and is used to filter text fragments that should not be identified as brand entities; When there are word conflicts among the precise brand dictionary, the fuzzy alias dictionary, and the exclusion word dictionary, the corresponding text fragments are filtered first based on the exclusion word dictionary.
7. The AI brand market share analysis method according to claim 5, characterized in that, The identification confidence level is used to determine the threshold and process low-confidence entities for the brand candidate entities, specifically including: Based on labeled AI response samples, several candidate confidence thresholds are set; Based on each candidate confidence threshold, the brand candidate entities are determined to be brand entities and non-brand entities, and the precision and recall corresponding to each candidate confidence threshold are calculated respectively. Based on the precision and the recall, calculate the harmonic evaluation value corresponding to each candidate confidence threshold, and determine the candidate confidence threshold corresponding to the maximum value of the harmonic evaluation value as the preset inclusion threshold; When the identification confidence of the brand candidate entity is greater than or equal to the preset inclusion threshold, the corresponding brand candidate entity is determined as a threshold-passing candidate entity. When the identification confidence of the brand candidate entity is lower than the preset inclusion threshold, the corresponding brand candidate entity is determined as a low-confidence candidate entity. Based on the ratio between the identification confidence level of the low-confidence candidate entity and the preset inclusion threshold, the statistical weight of the low-confidence candidate entity is proportionally reduced, or the low-confidence candidate entity is marked as pending confirmation.
8. The AI brand market share analysis method according to claim 1, characterized in that, The aforementioned brand candidate entities are merged using standard brand merging to obtain valid brand mentions bound to standard brand codes, including: The brand candidate entities are sequentially subjected to precise name matching, product name attribution matching, official website domain association, and enterprise entity association to obtain the brand merging result; When the brand merging result points to a unique standard brand code, the corresponding brand candidate entity is bound to the unique standard brand code; When product names are duplicated across brands or the relationship between brand entities is unclear, conflict resolution is carried out based on the context that has confirmed the co-occurrence relationship between the brands. When conflict resolution fails, the corresponding brand candidate entity will be marked as pending confirmation. Brand candidate entities that are bound to a unique standard brand code, appear in the valid content area of the response text, and whose identification confidence level meets the preset inclusion threshold are identified as the valid brand mentions.
9. The AI brand market share analysis method according to claim 1, characterized in that, The classification record includes the main category code, the list of secondary category codes, the confidence score of each category, the user decision-making influence score, and the question weight; The question weights are determined based on the classification records, including: Generate category word variations based on user-configured product or service names; Based on the product category variants, expand the purchasing questions and generate word-of-mouth verification questions; the word-of-mouth verification questions include brand names; Keyword extraction and semantic classification were performed on the question texts to which the valid brands mentioned belonged, and the main category code, the list of secondary category codes, and the confidence level of each category were determined. The user decision-making influence score is calculated based on issue coverage, purchase intent strength, and position sensitivity. The weight of a question is determined based on the proportion of the user's decision-making influence score in all question categories.
10. The AI brand market share analysis method according to claim 1, characterized in that, The determination of the occurrence position, recommendation ranking, and position weight includes: By parsing the answer text to understand its structure, we can obtain the recommendation list area, the main text paragraph area, the supplementary explanation area, and the cited source area. When the effective brand mention is located in the recommendation list area, the recommendation position is determined according to the list arrangement order; When the effective brand mention is not located in the recommended list area, the location of the mention is determined based on the positional relationship between the paragraph containing the effective brand mention and the beginning, middle, or end of the answer text. The location weight is determined based on the user attention decay relationship corresponding to the location of occurrence. The position weight is determined according to the decreasing display priority in the answer text.
11. The AI brand market share analysis method according to claim 10, characterized in that, The display priorities are as follows: The first item on the recommended list is higher than the second and third items; the second and third items on the recommended list are higher than the first paragraph of the main text; the first paragraph of the main text is higher than the fourth item and subsequent items on the recommended list. Recommended items ranked fourth and thereafter are ranked higher than the middle section of the main text; the middle section of the main text is ranked higher than the end of the main text; the end of the main text is ranked higher than the supplementary explanation area; and the supplementary explanation area is ranked higher than the source citation area.
12. The AI brand market share analysis method according to claim 1, characterized in that, The AI brand market share includes general mention share, answer coverage share, question coverage share, location-weighted share, sentiment-weighted share, and overall AI brand share. The methods for determining the market share of AI brands include: Calculate the general mention share based on the number of valid brand mentions corresponding to each standard brand code; The answer coverage share is analyzed based on the number of answers that mention the corresponding standard brand code at least once; The issue coverage share is determined by the number of issues that have occurred with the corresponding standard brand code; Calculate the position-weighted share by utilizing the position weight of each valid brand mentioned under the corresponding standard brand code; Based on the emotional tendencies mentioned by each valid brand under the corresponding standard brand code, an emotional weighted share is generated; The overall AI brand share is obtained by weighting and fusing the share of general mentions, the share of answer coverage, the share of question coverage, the share of location weighting, and the share of sentiment weighting.
13. The AI brand market share analysis method according to claim 1, characterized in that, The results of the multi-brand comparative analysis include an overview of AI brand market share, multi-brand rankings, problem details, platform comparison, trend analysis, sentiment comparison, and shortcomings. The AI brand market share overview is used to display the overall AI brand share, general mention share, position-weighted share, and issue coverage share of this product and its competitors. The multi-brand ranking is sorted according to the comprehensive AI brand share, number of mentions, average ranking, and positive mention ratio; The question details are used to display the mention, ranking, market share, and answer snippets of this product and competitors' products under each question.
14. An AI-powered brand market share analysis system, characterized in that: include: The brand entity configuration module is used to generate analysis tasks based on user configuration, including configurations for this product, competitors, brand entity, and analysis scope. The question scope configuration module is used to read AI response records based on the analysis task and determine the set of responses to be analyzed. The brand entity recognition module is used to perform brand entity recognition on the answer text in the set of answers to be analyzed, and obtain brand candidate entities; the brand candidate entities include entity name, text position, context fragments and recognition confidence; The standard brand merging module is used to perform standard brand merging on the brand candidate entities based on the brand entity configuration, so as to obtain valid brand mentions bound to the standard brand code; The issue stratification module is used to generate categorization records based on the issue text associated with the valid brand mentions, and to determine issue weights based on the categorization records; By determining the text position of the effective brand mentions in the answer text, the occurrence position, recommendation position and position weight are determined, and the emotional tendency corresponding to the effective brand mentions is identified. The market share calculation module is used to calculate the AI brand market share of this product and competitors based on the effective brand mentions, the question weight, the position weight, and the sentiment tendency. The comparative analysis module is used to generate multi-brand comparative analysis results based on the market share of the AI brands.
15. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 13.
16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 13.