Merchant entity name extraction method and device

By using pre-trained Chinese big model and iterative optimization technology, the problems of low accuracy and low efficiency of merchant name extraction in ringtone recording data are solved, efficient and accurate merchant name extraction is achieved, and the robustness and automation of the model are improved.

CN120354853APending Publication Date: 2025-07-22BEIJING YULORE INNOVATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510497322.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

When extracting merchant names from ringtone recording data, the prior art has problems such as low accuracy, low efficiency, poor robustness, and excessive dependence on manual labeled data.

Method used

Pre-trained Chinese large models (such as BERT) are used to combine B-NAME/I-NAME fine labeling and iterative optimization mechanisms, and through word segmentation, destop words, semantic analysis and deep learning model training, a merchant name extraction model is built, and the model parameters are iteratively optimized through verification feedback.

Benefits of technology

It significantly improves the recognition rate and accuracy of merchant names, enhances the stability of the model in a noisy and diverse text environment, reduces the dependence on initial labeled data, and improves processing efficiency and automation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354853A_ABST
    Figure CN120354853A_ABST
Patent Text Reader

Abstract

The invention discloses a merchant entity name extraction method and device, and the method comprises the steps: obtaining polyphonic ringtone recording data, and converting the polyphonic ringtone recording data into a text content set; performing word segmentation processing and stop word removal processing on the text content set to obtain a keyword set, and labeling the keyword set to obtain a labeled corpus; training a pre-trained Chinese big model based on the labeled corpus to obtain a merchant name extraction model; inputting a to-be-processed text into the merchant name extraction model, and extracting merchant name entities to obtain a merchant name extraction result set; verifying the merchant name extraction result set to obtain optimization feedback information; and according to the optimization feedback information, adjusting parameters of the merchant name extraction model, and outputting the optimized merchant name extraction model for extracting the merchant entity name. According to the invention, the extraction accuracy and the automation degree are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, in particular to the technical fields of artificial intelligence and natural language processing. More specifically, the present invention relates to a method and device for extracting specific entity information from speech-to-text conversion using a deep learning model, especially a method and device for extracting merchant entity names. Background Art

[0002] With the wide application of the outbound call system in fields such as commercial promotion, customer service, and marketing, the resulting ringback tone recording data has become an important source of information. The ringback tone often contains key merchant entity name information, and accurately extracting these names is of great value for subsequent data analysis, customer portrait construction, market strategy formulation, etc.

[0003] In the prior art, methods for extracting merchant names from unstructured text (especially text converted from speech recognition) mainly include rule-based matching and traditional machine learning-based methods. Rule-based methods usually set specific keywords or text patterns (such as "Welcome to XX", "This is XX Company", etc.) to identify merchant names. Traditional machine learning-based methods, such as models like Conditional Random Field (CRF), attempt to identify named entities by learning text features.

[0004] However, these prior arts have significant limitations when processing ringback tone recording conversion text. First, the speech recognition (ASR) process is prone to introducing noise, errors, and uncertainties, resulting in low-quality text content. Rule-based methods are difficult to adapt to such changes and have low matching accuracy. Second, the expressions of merchant names are extremely diverse, lacking a unified format and specification. Simple pattern matching or shallow machine learning models are difficult to generalize effectively and have poor recognition effects. Third, traditional machine learning models often require a large amount of high-quality manually annotated data for training, with high annotation costs, and the models have poor robustness to ASR errors and colloquial and non-standard expressions. In addition, existing methods often have low processing efficiency and are difficult to meet the real-time or near-real-time processing requirements of large-scale outbound call ringback tone data.

[0005] Therefore, there is an urgent need to propose a new technical solution that can overcome the deficiencies of the prior art, achieve efficient and accurate extraction of merchant entity names from ringback tone text data with noise and diverse formats, while improving the degree of automation and reducing the dependence on a large amount of initial annotated data. Summary of the Invention

[0006] In view of this, the main purpose of the present invention is to provide a method and device for extracting merchant entity names, aiming to solve the problems of low accuracy, low efficiency, poor robustness, and dependence on a large amount of manual annotation when extracting merchant names from ringback tone recording data.

[0007] Technical solution: To achieve the above object, the present invention provides a method for extracting merchant entity names, including: obtaining ringback tone recording data, and converting the ringback tone recording data into a text content set; performing word segmentation processing and stop word removal processing on the text content set to obtain a keyword set, and performing annotation on the keyword set to obtain an annotated corpus; training a pre-trained Chinese large model based on the annotated corpus to obtain a merchant name extraction model; inputting a text to be processed into the merchant name extraction model to extract merchant name entities, obtaining a merchant name extraction result set; verifying the merchant name extraction result set to obtain optimization feedback information; adjusting the parameters of the merchant name extraction model according to the optimization feedback information, and outputting an optimized merchant name extraction model for extracting merchant entity names.

[0008] Further, the performing word segmentation processing and stop word removal processing on the text content set includes: using a word segmentation tool to process the text content set to obtain a word sequence; removing stop words from the word sequence to form the keyword set; using the keyword set to extract entity information to obtain a candidate name set; and using B-NAME and I-NAME tags to process the candidate name set to construct the annotated corpus.

[0009] Further, the training the pre-trained Chinese large model based on the annotated corpus includes: using a data partitioning algorithm to process the annotated corpus to obtain a training set and a test set; training the pre-trained Chinese large model using the training set and outputting an initial model; using the test set to evaluate the performance of the initial model to obtain evaluation metrics; and adjusting the parameters of the initial model according to the evaluation metrics to obtain the merchant name extraction model.

[0010] Further, before adjusting the parameters of the initial model according to the evaluation metrics, it further includes: adjusting the evaluation metrics, including: using the initial model to generate a label prediction result as output data; constructing a hybrid model based on a deep belief Markov network and POMDP; inputting the output data into the hybrid model to obtain a probability distribution sequence; performing state inference calculation on the probability distribution sequence through the hybrid model to obtain an adjusted evaluation metric; and using the adjusted evaluation metric to adjust the parameters of the initial model.

[0011] Further, before performing word segmentation on the text content set, it includes: processing the text content set using a temporal encoder to obtain a time series representation; performing data dimensionality reduction processing on the time series representation to output a compressed sequence; constructing a semantic analysis model and using the semantic analysis model to process the compressed sequence to obtain semantic features related to the merchant name; wherein, performing word segmentation on the text content set includes: guiding the word segmentation processing using the semantic features related to the merchant name.

[0012] Further, performing data dimensionality reduction processing on the time series representation includes: constructing a semantic importance calculation model and using the semantic importance calculation model to process the time series representation to obtain an importance distribution; calculating a differential sampling frequency of the time series representation according to the importance distribution, and outputting a target sampling parameter including a sampling rate and sampling point positions; using the target sampling parameter to perform downsampling on the time series representation to obtain a sampled sequence; dividing the sampled sequence into multiple semantic units, and performing semantic integrity detection on the multiple semantic units to output the compressed sequence.

[0013] Further, before training the pre-trained Chinese large model based on the labeled corpus, it includes: performing data distribution calculation on the labeled corpus to obtain the frequency characteristics of the merchant name entity categories; calculating the data imbalance degree in the labeled corpus using the frequency characteristics of the merchant name entity categories; generating a sample weight coefficient according to the data imbalance degree; constructing a training loss function for the pre-trained Chinese large model using the sample weight coefficient, and determining the parameter settings for model training.

[0014] Further, calculating the data imbalance degree in the labeled corpus using the frequency characteristics of the merchant name entity categories includes: classifying the frequency characteristics of the merchant name entity categories according to the distribution characteristics to form multiple evaluation units; obtaining a feature subset of the merchant name entity categories from the multiple evaluation units; calculating the gradient change amount of the feature subset and outputting the gradient optimization direction; setting the training loss function parameters according to the gradient optimization direction and the data imbalance degree to generate a training objective function; using the training objective function to perform parameter adaptive adjustment on the pre-trained Chinese large model to generate a sample weight coefficient inversely proportional to the data imbalance degree.

[0015] Further, the step of inputting the text to be processed into the merchant name extraction model includes: performing text standardization processing on the text to be processed to obtain a normalized text; inputting the normalized text into the merchant name extraction model to obtain a character probability distribution; using the character probability distribution to identify the entity boundary positions and output a position marking sequence; and extracting corresponding text segments from the normalized text according to the position marking sequence to generate the merchant name extraction result set.

[0016] Further, the step of verifying the merchant name extraction result set includes: evaluating the merchant name extraction result set using a verification data set to obtain an abnormal sample set; expanding the labeled corpus according to the abnormal sample set to form an enhanced training set; analyzing the performance bottleneck of the model based on the enhanced training set to determine the model structure parameters to be optimized and generating a model tuning scheme; executing the model tuning scheme to train the merchant name extraction model, and calculating the recognition rate and accuracy rate of the model on the verification data set to form optimization feedback information including performance indicators and error cases.

[0017] The present invention also provides an extraction device for merchant entity names, including: a data acquisition module for acquiring ringback tone recording data and converting the ringback tone recording data into a text content set; a data processing module for performing word segmentation processing and stop word removal processing on the text content set to obtain a keyword set and performing annotation based on the keyword set to obtain a labeled corpus; a model training module for training a pre-trained large Chinese model based on the labeled corpus to obtain a merchant name extraction model; a name extraction module for inputting the text to be processed into the merchant name extraction model to extract merchant name entities and obtain a merchant name extraction result set; a verification module for verifying the merchant name extraction result set to obtain optimization feedback information; and an iterative optimization module for adjusting the parameters of the merchant name extraction model according to the optimization feedback information and outputting an optimized merchant name extraction model for extracting merchant entity names.

[0018] Compared with the prior art, the present invention has the following beneficial effects: Improved extraction accuracy: By adopting a pre-trained large Chinese model (such as BERT) and combining with fine-tuning for specific tasks, as well as introducing B-NAME / I-NAME fine annotation and an iterative optimization mechanism, it is possible to more accurately identify and extract various forms of merchant names, significantly improving the recognition rate and accuracy rate (for example, in the embodiment, the recognition rate can reach 99.3% and the accuracy rate can reach 96.1%).

[0019] Enhanced robustness: The deep learning model has better fault tolerance for the noise brought by speech recognition and the diversity of text formats; combined with extended technologies such as deep belief Markov network, POMDP, semantic-aware compression, and debiased gradient extrapolation, the stability and generalization ability of the model in complex, uncertain, and biased data environments are further enhanced.

[0020] Improved processing efficiency and automation: The entire process forms a closed loop from data acquisition, processing, model training, extraction to optimization, reducing manual intervention and achieving a high degree of automation; technologies such as semantic-aware compression help improve the speed of large-scale data processing.

[0021] Reduced dependence on initial labeled data: Utilizing the transfer learning ability of the pre-trained model, training can be initiated with relatively less labeled data, and the performance can be gradually improved through iterative optimization and active learning (such as expanding the corpus according to validation feedback). BRIEF DESCRIPTION OF THE DRAWINGS

[0022] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required for use in the embodiments will be briefly introduced below. The accompanying drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure. It should be understood that the following drawings only illustrate certain embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0023] Figure 1 It is a schematic flow chart of a method for extracting merchant entity names in an embodiment of the present invention; Figure 2 It is a schematic structural diagram of a device for extracting merchant entity names in an embodiment of the present invention; Figure 3 It is a schematic flow chart of a method for semantic-aware data compression in an embodiment of the present invention; Figure 4 It is a schematic flow chart of the model training process for data imbalance processing in an embodiment of the present invention; Figure 5 It is a schematic flow chart of the result set of merchant name extraction provided in an embodiment of the present invention; Figure 6 It is a schematic flow chart of obtaining optimization feedback information in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part rather than all of the embodiments of the present disclosure. Components of the embodiments of the present disclosure described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present disclosure provided in the accompanying drawings is not intended to limit the scope of the present disclosure claimed, but merely represents selected embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative efforts fall within the scope of protection of the present disclosure.

[0025] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0026] As used herein, the term "and / or" merely describes an association relationship and indicates that three relationships may exist. For example, A and / or B may represent three cases: A exists alone, both A and B exist simultaneously, and B exists alone. Additionally, the term "at least one" as used herein means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent including any one or more elements selected from the set composed of A, B, and C.

[0027] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0028] The following further elaborates in detail on the embodiments of the present invention with reference to the accompanying drawings.

[0029] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this invention belongs. The terms used in the specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0031] Embodiment 1 This embodiment provides a method for extracting merchant entity names. Referring to Figure 1 as shown, the process may include the following steps: S1: Obtain the ringback tone recording data and convert the ringback tone recording data into a set of text contents.

[0032] Specifically, first, batch collect ringback tone recording samples containing potential merchant information through the interface or data collection module of the outbound call system to form an original recording data set.

[0033] To ensure the quality of subsequent processing, it is necessary to evaluate the quality of the collected recordings, such as performing signal-to-noise ratio analysis, volume normalization, removal of silent segments, identification of valid speech segments, etc., and screen out a valid recording set with qualified quality.

[0034] Then, input the valid recording set into a professional Automatic Speech Recognition (ASR) engine, such as using technologies like Google Speech-to-Text, Baidu Speech Recognition API, iFlytek, or the open-source Sphinx, etc., to convert the speech signal into the initial text content.

[0035] Since errors and non-text characters may be introduced during the ASR process, it is necessary to perform basic text cleaning and normalization processing on the converted initial text, such as removing special symbols (such as *, #, etc.), correcting common recognition errors (such as homophone replacement), unifying the text format (such as full-width to half-width), etc., and finally output a relatively clean and normalized set of text contents as the processing object for subsequent steps.

[0036] S2: Perform word segmentation processing and stop word removal processing on the set of text contents to obtain a set of keywords, and label the set of keywords to obtain a labeled corpus.

[0037] First, perform Chinese word segmentation processing on the set of normalized text contents output in step S1. Mature Chinese word segmentation tools, such as jieba, HanLP, PKUSeg, etc., can be used to cut the continuous text into a discrete sequence of words to form a word segmentation result set. Word segmentation is the basis for understanding the semantics of Chinese text.

[0038] Next, based on the word segmentation result set, stop words are removed. Stop words refer to words that appear frequently in the text but do not contribute much to the expression of the core semantics (the merchant name in this scenario), such as "的", "了", "是", "欢迎光临", "你好", etc. By loading a predefined stop word list (which can be a general stop word list or a word list customized based on the characteristics of the field), these stop words are filtered out from the word sequence to obtain a keyword set that better reflects the text theme and key information.

[0039] Then, the processed keyword set is used to extract potential business name entity information to form a candidate name set. This step can be done by identifying noun phrases through part-of-speech tagging, combining heuristic rules such as position features (such as whether it appears after a specific sentence structure), or preliminary pattern matching to screen possible business name fragments.

[0040] Finally, the candidate name set is manually annotated to build a high-quality annotated corpus. This is the key to supervised learning model training. Professional data annotation platforms (such as doccano, or others such as brat, Label Studio, etc.) can be used to assist in the annotation work. The annotator determines the real business name in the text based on the context and annotates it accurately.

[0041] This embodiment preferably adopts the B-NAME (Begin of Name) and I-NAME (Inside of Name) marking methods in the BIESO or BIO marking system. For example, for "XX restaurant", it is marked as "XX / B-NAME restaurant / I-NAME". This marking method can clearly indicate the entity boundaries and help the model learn the beginning and internal structure of the name. After the marking is completed, a marked corpus containing the original text (or processed text) and the corresponding label sequence is formed for subsequent model training.

[0042] In one embodiment, in step S2, the word segmentation and stop word removal processing of the text content set specifically includes: S2.1: Use a word segmentation tool to process the text content set to obtain a word sequence.

[0043] As mentioned earlier, use tools such as jieba for word segmentation.

[0044] S2.2: Remove stop words from the word sequence to form the keyword set.

[0045] As mentioned before, word sequences are filtered based on a stop word list.

[0046] S2.3: Extract entity information using the keyword set to obtain a candidate name set.

[0047] As mentioned above, screen candidate names through features such as part of speech and position.

[0048] S2.4: Process the candidate name set using B-NAME and I-NAME tags to construct the labeled corpus.

[0049] As mentioned above, use a labeling platform and B-NAME / I-NAME tagging method for manual annotation.

[0050] The core processes of text data preprocessing and corpus annotation include word segmentation, stop word removal, candidate name extraction, and manual annotation.

[0051] First, use mature Chinese word segmentation tools, such as jieba, HanLP, etc., to process the normalized text content set, and segment the continuous text string into a sequence of words with independent meanings, which is the basis for subsequent natural language processing tasks.

[0052] Next, in order to reduce noise and focus on key information, stop word removal processing is required. By loading a predefined or customized stop word dictionary (including words such as "de", "le", "shi", "nin hao" that are common but contribute little to identifying merchant names), remove these stop words from the word sequence obtained in the previous step to form a more refined keyword set.

[0053] Subsequently, based on this keyword set, initially extract the potential entity information in the text that may be the merchant name by applying part-of-speech tagging to identify noun phrases, analyzing the position characteristics of words in sentences, or using simple pattern matching, etc., to form a candidate name set.

[0054] Finally, and most importantly, is to accurately manually annotate these candidate name sets to construct high-quality training data. This process can be carried out with the help of a professional data annotation platform (such as doccano). The annotator accurately identifies the real merchant names in the text according to the context semantics, and adopts a specific annotation system, preferably using tags such as B-NAME (indicating the starting position of the name) and I-NAME (indicating the non-starting position of the name), to annotate each character or word that makes up the merchant name.

[0055] For example, "XX Restaurant" will be annotated as "XX / B-NAME Restaurant / I-NAME". This word-by-word or character-by-character boundary annotation method helps the model accurately learn the structure and boundary information of the merchant name. After completion of the annotation, a labeled corpus containing the original text (or its processed form) and the corresponding tag sequence is obtained as the input for subsequent deep learning model training.

[0056] Optionally, as Figure 3 shown, before performing word segmentation on the text content set in step S2, the method further includes: A1: Process the text content set using a temporal encoder to obtain a time series representation.

[0057] Process the text content set (essentially a temporal sequence of words or characters) using a temporal encoder (such as the encoding layer of RNN, LSTM, or Transformer) to obtain its time series representation, which can capture the sequential relationships and context dependencies between words.

[0058] A2: Perform data dimensionality reduction on the time series representation and output a compressed sequence.

[0059] Perform data dimensionality reduction on this time series representation to retain key semantic information and remove redundancy. Specifically: A2.1: Construct a semantic importance calculation model and use the semantic importance calculation model to process the time series representation to obtain an importance distribution.

[0060] A2.2: According to the importance distribution, calculate the differential sampling frequency of the time series representation and output target sampling parameters including the sampling rate and the positions of sampling points.

[0061] A2.3: Use the target sampling parameters to downsample the time series representation to obtain a sampled sequence.

[0062] A2.4: Divide the sampled sequence into multiple semantic units and perform semantic integrity detection on the multiple semantic units, and output the compressed sequence.

[0063] This dimensionality reduction process can be achieved by constructing a semantic importance calculation model, analyzing the semantic importance of each part in the time series representation for identifying the merchant name, and obtaining the importance distribution. According to this importance distribution, calculate the differential sampling frequency, that is, maintain a high sampling rate for semantically important parts and reduce the sampling rate for less important parts, and output target sampling parameters including the sampling rate and the positions of sampling points. Use this parameter to downsample the time series representation to obtain a sampled sequence. Further, the sampled sequence can be divided into multiple semantic units and semantic integrity detection can be performed to ensure that the compressed sequence is still relatively complete semantically, and finally output the compressed sequence.

[0064] A3: Construct a semantic analysis model and use the semantic analysis model to process the compressed sequence to obtain semantic features related to the merchant name.

[0065] Build a semantic analysis model to process the compressed sequence and extract semantic features related to the merchant name.

[0066] Among them, in step S2, the tokenization of the text content set includes: using the semantic features related to the merchant name to guide the tokenization process.

[0067] These extracted semantic features can be used to guide subsequent tokenization. For example, splitting is preferably performed at the boundaries of semantic units, or the dictionary or model of the tokenizer is adjusted using semantic features to make it more likely to recognize potential merchant names as a single word. This approach helps reduce the data volume while retaining key information, improves processing efficiency, and may enhance the tokenization quality.

[0068] In an alternative implementation, before tokenizing the text content set, a semantic-aware temporal data processing and compression technique can be introduced to optimize the data representation and improve subsequent processing efficiency. This process first uses a temporal encoder, such as the encoding part of a recurrent neural network (RNN), long short-term memory network (LSTM), or Transformer architecture, to process the input text content set, capturing the sequential dependencies between words or characters to obtain a temporal sequence representation of the text. Subsequently, to reduce data redundancy while retaining key semantic information, data dimensionality reduction processing is required on the generated temporal sequence representation, and a compressed sequence is output.

[0069] Specifically, this data dimensionality reduction processing can be further refined as follows: build a semantic importance calculation model that is responsible for analyzing the semantic contribution of different parts of the temporal sequence representation to the core task of identifying the merchant name, and then obtain an importance distribution map. Based on this importance distribution, the system can calculate the differential sampling frequency required for the temporal sequence representation, that is, a higher sampling rate is used for regions with high semantic importance (such as parts that may contain the merchant name), while a lower sampling rate is used for regions with low semantic importance, and finally, the target sampling parameters including the specific sampling rate and sampling point positions are output. Using these target sampling parameters, a downsampling operation is performed on the original temporal sequence representation to obtain a sampled sequence with reduced data volume but relatively high information density.

[0070] To ensure semantic coherence, the sampled sequence can be further divided into multiple semantic units, and semantic integrity detection is performed on these units, such as checking the semantic consistency within the unit or whether it forms a relatively complete meaning segment, and finally, a verified compressed sequence is output.

[0071] After compression, a semantic analysis model is constructed and used to process the compressed sequence with the aim of extracting semantic features highly relevant to merchant name recognition. These extracted semantic features can then effectively guide subsequent word segmentation processing steps. For example, by inputting the semantic features as additional information into the word segmenter or using the features to adjust the judgment logic of word segmentation boundaries, making it more inclined to segment potential merchant names as a whole, thereby improving processing efficiency while potentially enhancing the accuracy of word segmentation and laying a better foundation for subsequent model training.

[0072] S3: Train a pre-trained large Chinese model based on the annotated corpus to obtain a merchant name extraction model.

[0073] First, divide the annotated corpus produced in step S2. Usually, it is divided into a training set and a test set (and sometimes a validation set) according to a certain ratio (8:2). The training set is used to train model parameters, and the test set is used to evaluate the generalization ability of the model on unseen data.

[0074] Then, select a pre-trained large Chinese language model as the base model, such as "bert-base-chinese". On this basis, construct a neural network architecture suitable for the named entity recognition (NER) task, usually by connecting one or more linear layers (classifiers) after its output layer.

[0075] Next, use the training set data to perform fine-tuning training on this model to output an initial model. Key training parameters need to be configured during the training process, such as batch size (batch_size = 32), learning rate (lr = 3e-5), number of training epochs (epochs = 3), and select an optimizer (such as AdamW) and possible regularization strategies (such as weight decay, gradient clipping max_grad_norm = 1.0).

[0076] The goal of training is to minimize the loss function between the predicted labels and the true labels. During or after the training process, use the test set to evaluate the performance of the initial model and calculate evaluation metrics (such as accuracy, recall, F1 value). Finally, adjust the model parameters or training strategies according to the evaluation metrics (possibly combined with the performance on the validation set), such as using the early stopping (EarlyStopping) strategy to prevent overfitting, and finally select the model version with the best performance as the merchant name extraction model.

[0077] In one embodiment, in step S3, the training of the pre-trained large Chinese model based on the annotated corpus specifically includes: S3.1: Process the annotated corpus using a data partitioning algorithm to obtain a training set and a test set.

[0078] As described above, divide the dataset proportionally.

[0079] S3.2: Train the pre-trained Chinese large model using the training set and output the initial model.

[0080] As described above, use the training set for model fine-tuning.

[0081] S3.3: Evaluate the performance of the initial model using the test set to obtain evaluation metrics.

[0082] As described above, use the test set to calculate performance metrics.

[0083] S3.4: Adjust the parameters of the initial model according to the evaluation metrics to obtain the merchant name extraction model.

[0084] As described above, select the optimal model or adjustment strategy based on the evaluation results.

[0085] The typical process of training a merchant name extraction model based on an annotated corpus includes key links such as data division, model training, performance evaluation, and parameter adjustment.

[0086] First, the prepared annotated corpus needs to be divided into a training set and a test set (and sometimes a validation set) using a data division algorithm according to a predetermined rule (for example, the commonly used ratio of 80% for training and 20% for testing), ensuring that the data distribution between different sets is as consistent as possible. The training set is used for model learning, and the test set is used to evaluate the generalization ability of the model.

[0087] Next, select a powerful pre-trained Chinese language model (such as "bert-base-chinese" mentioned in the technical disclosure) as the basic architecture, and build a specific network layer for the named entity recognition (NER) task on it. Usually, one or more fully connected layers are added after the output layer of BERT as a classifier, which is responsible for predicting the label (such as B-NAME, I-NAME, O) of each input unit (token).

[0088] Then, use the divided training set data to fine-tune the constructed model. During this process, training hyperparameters need to be carefully set, such as batch size, learning rate, number of training epochs, and select a suitable optimizer (such as AdamW) and possible regularization techniques (such as weight decay, gradient clipping) to guide the update of model parameters, with the goal of minimizing the loss between the model prediction output and the true labels of the training set.

[0089] During or after completing the training, it is necessary to use the test set (or validation set) to evaluate the performance of the current model (initial model), calculate a series of key evaluation metrics, such as Precision, Recall, and F1-score, which can quantify the performance of the model in the task of identifying merchant names. Finally, based on the results of these evaluation metrics, the parameters of the model are finally adjusted or selected.

[0090] For example, the Early Stopping strategy can be adopted, that is, when the performance of the model on the validation set no longer improves, the training is terminated in advance to prevent overfitting; or the best model checkpoint is selected according to the performance on the validation set. Through this series of steps, an optimized deep learning model that can be used for the actual merchant name extraction task is finally obtained.

[0091] Optionally, before adjusting the parameters of the initial model according to the evaluation metrics in step S3.4, the method further includes: S3.4.1: Adjust the evaluation metrics, including: using the initial model to generate label prediction results as output data.

[0092] To handle the uncertainty brought by speech recognition, first use the initial model to generate preliminary label prediction results (such as probability distribution).

[0093] S3.4.2: Construct a hybrid model based on the deep belief Markov network and POMDP.

[0094] Construct a hybrid model that combines DCMN (capturing long-distance dependencies and confidence) and POMDP (handling uncertainty decisions).

[0095] S3.4.3: Input the output data into the hybrid model to obtain a sequence of probability distributions.

[0096] Input the output of the initial model into the hybrid model to obtain a sequence of probability distributions processed by DCMN.

[0097] S3.4.4: Perform state inference calculation on the sequence of probability distributions through the hybrid model to obtain adjusted evaluation metrics.

[0098] Use the POMDP inference framework to perform state inference on the probability sequence, find the optimal label sequence under uncertainty, so as to obtain adjusted evaluation metrics that can better reflect the model's ability under uncertainty, or directly obtain a better label sequence.

[0099] S3.4.5: Use the adjusted evaluation metrics to adjust the parameters of the initial model.

[0100] Use these adjusted metrics to guide the final adjustment of the model parameters, or use this hybrid model as the final model to enhance robustness.

[0101] To further improve the robustness of the model when processing speech recognition texts with noise and uncertainty, a hybrid inference framework combining Deep Confidence Markov Network (DCMN) and Partially Observable Markov Decision Process (POMDP) can be introduced before the basic model training is completed and the final parameter adjustment is carried out. This process first uses the trained initial model to make a preliminary prediction on the data (such as the validation set or test set), generating the label probability distribution or the preliminary label sequence for each token, and these prediction results will be used as the input data for the subsequent hybrid model.

[0102] Next, construct a hybrid model architecture that integrates DCMN and POMDP. DCMN is good at capturing long-distance dependencies in sequential data and combining confidence information of predictions, while POMDP is a powerful mathematical framework for making optimal decisions in the case of incomplete information (partially observable).

[0103] Input the output of the initial model (such as the label probability sequence) into this hybrid model. The DCMN module will first process it, considering sequential dependencies and confidence, and output a refined probability distribution sequence.

[0104] Subsequently, the POMDP inference framework intervenes, modeling the label prediction problem as a sequential decision-making process, where the true label is regarded as a partially observable state. POMDP performs state inference calculations under the uncertain probability distribution based on historical observations and the current context, aiming to find the globally optimal label sequence.

[0105] Through the inference process of this hybrid model, a set of adjusted evaluation metrics (such as performance metrics calculated based on the final decision of POMDP) or a corrected and more confident label sequence can be obtained.

[0106] Finally, use these adjusted evaluation metrics that can better reflect the true capabilities of the model in an uncertain environment to guide the final selection and optimization of the initial model parameters. Or in some implementations, use the entire BERT+DCMN+POMDP hybrid architecture as the final deployed merchant name extraction model, thus effectively improving the accuracy and stability of the model when facing complex inputs.

[0107] The hybrid inference framework of the Deep Belief Markov Network (DCMN) and the Partially Observable Markov Decision Process (POMDP) aims to address the lack of robustness that may occur in standard sequence labeling models (such as those that only use BERT+CRF or BERT+Softmax) when dealing with uncertain inputs (especially texts with noise and ambiguity originating from ASR). Standard models often make decisions based on local information and maximum probability, and are easily interfered by local noise, leading to errors in the global label sequence. This hybrid framework alleviates this problem by introducing more complex probability modeling and decision-making mechanisms.

[0108] First, the Deep Belief Markov Network (DCMN) is used to deeply process and model the confidence of the sequence information initially predicted by the base model (such as BERT). BERT can provide powerful context representations, but the label probabilities it outputs for each token may themselves be uncertain.

[0109] Instead of simply accepting these probabilities, DCMN treats them as signals with confidence information. It uses a neural network structure (possibly an RNN or a Transformer variant) to capture long-range dependencies in the label sequence and explicitly models the confidence of each prediction. This means that DCMN can learn typical transition patterns between labels (for example, "B-NAME" is usually followed by "I-NAME" or "O", but rarely directly by another "B-NAME"), and can judge whether the model's prediction at a certain position is "certain" or "hesitant".

[0110] By combining context dependence and confidence evaluation, DCMN can output a more refined probability distribution sequence or feature representation that better reflects the overall structural rationality of the sequence and the reliability of each part of the prediction.

[0111] Next, the Partially Observable Markov Decision Process (POMDP) elevates the sequence labeling problem to a more advanced decision-making level. It treats the true label sequence as a hidden (partially observable) state sequence, and all that the model can directly obtain are the observations (i.e., the features / probabilities from BERT or processed by DCMN).

[0112] The goal of POMDP is to find the optimal strategy to infer the most likely hidden state sequence (i.e., the most accurate label sequence) under this uncertainty. It needs to define the state space (all possible label sequences), the observation space (the features / probabilities output by the model), the state transition probability (the likelihood of a label transitioning to the next label, which may depend on the context), the observation probability (the likelihood of observing a specific feature / probability under a certain true label), and a reward function (a criterion for measuring the accuracy of the final label sequence).

[0113] Through complex inference algorithms such as particle filter - based or point - based variants of value iteration, POMDP can calculate the label sequence with the maximum expected cumulative reward considering all possible paths and probabilities. This means that POMDP does not simply select the most likely label at each position, but rather searches for the globally optimal label sequence that is the most coherent, consistent with the probability model, and overall optimal.

[0114] When deploying the entire BERT + DCMN + POMDP hybrid architecture as a whole, its workflow is as follows: The input text to be processed first obtains the initial context - aware representation and preliminary label probabilities through BERT; this information is then fed into DCMN for sequence - dependency modeling and confidence evaluation, and refined probabilities or features are output; finally, POMDP receives the information from DCMN as observations, performs powerful probabilistic inference and decision - making processes, and outputs the final and optimal merchant name label sequence.

[0115] This multi - stage, step - by - step architecture for dealing with uncertainty makes the model more robust when facing complex inputs (such as typos, homophone substitutions, incomplete sentences caused by ASR errors, or ambiguous and vague merchant names themselves).

[0116] For example, even if ASR recognizes "Starlight Fast Food" as "Starlight Fast Can", BERT may give a low confidence score for the character "Can". DCMN will capture this low confidence score and the common pattern that a noun component follows "Fast". When making a decision, POMDP will comprehensively consider the observed low - confidence features, possible true labels (the possibility of "Food" still exists and it collocates more naturally with "Fast"), and label transition probabilities, thus having a higher probability of inferring the correct label sequence or a more reasonable approximate sequence, rather than being completely misled by the incorrect observations. Therefore, this hybrid architecture can effectively improve the accuracy and stability of the model in a noisy real - world data environment.

[0117] Optionally, as Figure 4 shown, before training the pre - trained large - scale Chinese model based on the labeled corpus in step S3, the method further includes: B1: Perform data distribution calculation on the labeled corpus to obtain the frequency features of merchant name entity categories.

[0118] There may be a data imbalance problem in the ringtone data. First, perform data distribution analysis on the labeled corpus to obtain the frequency features of each category (such as different types of merchant names or simple yes / no).

[0119] B2: Calculate the data imbalance degree in the labeled corpus by using the frequency characteristics of the merchant name entity categories.

[0120] Calculate the data imbalance degree based on the frequency characteristics, such as through metrics like class ratio or Gini coefficient.

[0121] B3: Generate sample weight coefficients according to the data imbalance degree.

[0122] Generate different weights for samples of different classes. Usually, the weights of minority class samples are higher.

[0123] B4: Use the sample weight coefficients to construct the training loss function of the pre-trained Chinese large model and determine the parameter settings for model training.

[0124] Integrate the sample weight coefficients into the loss function (such as weighted cross-entropy) so that the model pays more attention to the minority class during training to mitigate the impact of data imbalance.

[0125] Optionally, in step B2, the calculating the data imbalance degree in the labeled corpus by using the frequency characteristics of the merchant name entity categories may specifically include: B2.1: Classify the frequency characteristics of the merchant name entity categories according to the distribution characteristics to form multiple evaluation units.

[0126] B2.2: Obtain the feature subsets of the merchant name entity categories from the multiple evaluation units.

[0127] B2.3: Calculate the gradient change amount of the feature subsets and output the gradient optimization direction.

[0128] B2.4: Set the parameters of the training loss function according to the gradient optimization direction and the data imbalance degree, and generate the training objective function.

[0129] B2.5: Use the training objective function to adaptively adjust the parameters of the pre-trained Chinese large model to generate sample weight coefficients that are inversely proportional to the data imbalance degree.

[0130] This more refined method identifies the sources of bias by analyzing the class feature distribution and calculating the gradient change, and accordingly (such as using gradient extrapolation techniques) sets the parameters of the training objective function, adaptively generates more effective sample weights, and conducts debiased representation learning.

[0131] Considering that the labeled merchant name samples in the actual ringtone data may be relatively sparse, or the frequencies of certain types of merchant names (such as specific industries or lengths) are much higher than those of other types, which easily leads to unbalanced training data, thereby affecting the generalization ability and fairness of the model.

[0132] To address this challenge, a data imbalance handling strategy can be implemented before the formal start of model training. This strategy first requires a detailed calculation of the data distribution of the annotated corpus, and statistical analysis of the sample quantity and frequency characteristics of different merchant name entity categories (the categories here can be defined according to dimensions such as length, industry, or simply divided into the existence / non-existence of merchant names, etc.). Based on these statistically obtained frequency characteristics, the data imbalance degree of the entire annotated corpus is then calculated, which can be quantified by calculating the sample ratio between categories, information entropy, Gini coefficient, or other relevant metrics.

[0133] According to the quantified imbalance degree, corresponding sample weight coefficients are generated for samples of different categories in the dataset. The basic principle is to assign higher weights to categories with fewer samples to increase their influence during the training process. Finally, these generated sample weight coefficients are integrated into the training loss function of the pre-trained Chinese large model. For example, using weighted cross-entropy loss makes the contribution of each sample to the overall loss proportional to its weight. By adjusting the loss function in this way, the model can be effectively guided to pay more attention to those underrepresented minority class samples during training, thereby alleviating the negative impact brought by data imbalance and determining the parameter settings for model training (especially the specific form of the loss function).

[0134] Furthermore, the process of calculating the data imbalance degree and generating sample weights can adopt a more refined, gradient-based adaptive adjustment method to achieve more effective debiased representation learning.

[0135] This method first classifies the previously obtained merchant name entity category frequency characteristics according to their distribution characteristics (for example, whether they belong to the tail of the long-tailed distribution) to form multiple different evaluation units. Then, representative feature subsets for each category are selected from these evaluation units.

[0136] In the early stage of model training or through pre-analysis, calculate the gradient change amount of the model when processing these feature subsets. Analyzing the gradient information helps identify which features or categories are the main sources of learning bias, and accordingly output the direction of gradient optimization, that is, how the model parameters should be adjusted to reduce bias.

[0137] Then, combining this gradient optimization direction and the previously calculated data imbalance degree information, more precisely set the parameters of the training loss function, which may include dynamically adjusting weights, introducing regularization terms, or using gradient extrapolation techniques to predict the optimization path on ideal balanced data, thereby generating a more effective training objective function.

[0138] Finally, the optimized training objective function is used to adaptively adjust the parameters of the pre-trained Chinese large model, which means that the weights of the samples may change dynamically during training, and finally generate sample weight coefficients that are inversely proportional or have a more complex relationship with the data imbalance degree and can effectively suppress bias. This strategy based on gradient information and adaptive adjustment can more deeply solve the problem of internal data bias and prompt the model to learn more fair and robust feature representations.

[0139] S4: Input the text to be processed into the merchant name extraction model, extract merchant name entities, and obtain a merchant name extraction result set.

[0140] When the trained model is needed for actual merchant name extraction, that is, in the inference stage of processing the text to be processed, the process usually includes text normalization, model prediction, boundary recognition, and entity extraction.

[0141] First, perform text normalization on the input original text, which may include steps similar to the preprocessing of training data, such as removing irrelevant characters, unifying formats (such as full-width to half-width), processing special symbols, etc. The purpose is to obtain a normalized text with a standardized format that meets the input requirements of the model.

[0142] Next, input this normalized text into the merchant name extraction model that has been trained and deployed. After receiving the input, the model performs forward propagation calculations. For each basic unit (token, usually a character or a word) in the text sequence, the model will output the probability distribution of its belonging to each predefined label category (such as B-NAME, I-NAME, O-Outside).

[0143] Based on this probability distribution output by the model, the next step is to identify the entity boundary positions. The common approach is to adopt a greedy strategy, select the label with the highest probability for each token, thus forming a preliminary label sequence. By parsing this label sequence, the tokens marked as B-NAME (start of name) and all consecutive tokens marked as I-NAME (inside name) following it can be located. These consecutive B-I marks jointly define the scope of a merchant name entity, and finally output a position marking sequence that can represent the entity boundary.

[0144] Finally, according to this exact sequence of position markers, extract the corresponding text fragments from the input normalized text. For example, if the sequence indicates that the 5th to 8th characters form a merchant name, extract this substring. Collect all the text fragments identified and extracted in this way to form the final set of merchant name extraction results. In some cases, some post-processing operations may be required, such as removing entities that are too short or illogical, merging adjacent entities of the same type, etc., to improve the quality of the results.

[0145] Among them, as Figure 5 shown, step S4 specifically includes: S4.1: Perform text normalization processing on the text to be processed to obtain normalized text.

[0146] Before inputting into the model, perform normalization processing such as cleaning and format unification on the text to be processed to ensure that it meets the model input requirements.

[0147] S4.2: Input the normalized text into the merchant name extraction model to obtain a character probability distribution.

[0148] Input the processed text into the model, and the model performs forward calculation to output the probability that each token (word / character) belongs to each label (B-NAME, I-NAME, O, etc.).

[0149] S4.3: Use the character probability distribution to identify the entity boundary positions and output a sequence of position markers.

[0150] Select the label of each token according to the highest probability to form a label sequence, and identify the start and end boundaries of the merchant name entity according to the B-NAME and I-NAME markers.

[0151] S4.4: Extract the corresponding text fragments from the normalized text according to the position marker sequence to generate the set of merchant name extraction results.

[0152] According to the identified boundary positions, extract the corresponding strings from the original text as merchant name entities, and collect all the extracted entities to form a result set. Post-processing may be required.

[0153] S5: Verify the set of merchant name extraction results to obtain optimization feedback information.

[0154] Such as Figure 6 shown, step S5 specifically includes: S5.1: Use the validation dataset to evaluate the set of merchant name extraction results to obtain a set of abnormal samples.

[0155] Evaluate the output results of the current model using independent validation data, calculate performance metrics, and identify samples with prediction errors or omissions to form an abnormal sample set.

[0156] S5.2: Expand the labeled corpus according to the abnormal sample set to form an enhanced training set.

[0157] Add the identified abnormal samples (especially hard cases and edge cases) to the labeled corpus, perform annotation or correct annotation to form an enhanced training set.

[0158] S5.3: Analyze the performance bottleneck of the model based on the enhanced training set, determine the model structure parameters to be optimized, and generate a model tuning plan.

[0159] Combine error analysis and enhanced data to analyze the reasons for the poor performance of the model, determine whether it is a problem of model structure, parameter settings or training strategy, and formulate a corresponding model tuning plan.

[0160] S5.4: Execute the model tuning plan to train the merchant name extraction model, and calculate the recognition rate and accuracy of the model on the validation data set to form optimization feedback information including performance metrics and error cases.

[0161] According to the tuning plan, use the enhanced training set to retrain or continue training the model. Then evaluate the performance of the new model on the validation set again, record the performance metrics and error cases, and form complete optimization feedback information.

[0162] To ensure the continuous stability and improvement of the model performance, a closed-loop process of verification and iterative optimization needs to be established. This process starts with the verification and evaluation of the current output results of the model.

[0163] Specifically, use a reserved validation data set independent of the training data to evaluate the extraction result set generated by the merchant name extraction model. By comparing the model prediction results with the true labels on the validation set, calculate key performance metrics such as accuracy, recall rate, and F1 value, and at the same time identify and collect samples with model prediction errors or failures to identify, forming an abnormal sample set, which reflects the current weaknesses of the model.

[0164] Next, according to the problems revealed by analyzing these abnormal sample sets, strategically expand the original labeled corpus. In particular, these "hard cases" that the model handles poorly and "edge cases" at the decision boundary should be added to the corpus and accurately annotated, so as to form an enhanced training set with higher data quality and more comprehensive coverage scenarios.

[0165] Based on an in-depth analysis of abnormal samples and the enhanced training set, further diagnose the bottleneck of the model performance, and determine whether the problem stems from the model structure design, insufficient feature representation ability, or the training strategy needs to be improved. Accordingly, determine the specific model structure parameters (such as adjusting the number of network layers, replacing the activation function) or training hyperparameters (such as the learning rate scheduling strategy, optimizer selection) that need to be optimized, and finally form a clear model tuning plan.

[0166] Finally, strictly execute this model tuning plan, and retrain or incrementally train the merchant name extraction model using the enhanced training set. After the training is completed, it is necessary to comprehensively evaluate the performance of the new version model on the validation dataset again, calculate and record metrics such as its recognition rate and accuracy, and organize these performance data together with the specific error case analysis into an optimization feedback message. This feedback message is not only used to measure the effect of this optimization iteration, but also provides a decision-making basis and direction for the next possible optimization.

[0167] S6: Adjust the parameters of the merchant name extraction model according to the optimization feedback message, and output an optimized merchant name extraction model for extracting merchant entity names.

[0168] Based on the optimization feedback message generated in step S5, determine whether the newly trained model is better than the old model. If the performance is significantly improved (such as the recognition rate and accuracy are increased), then deploy this model with adjusted and optimized parameters to the actual application to replace the old model. If the performance does not improve or decreases, it may be necessary to re-analyze and formulate a new tuning strategy. Through this continuous "verification - feedback - adjustment - optimization" cycle, the model performance can be continuously improved.

[0169] It should be noted that the above steps are only divided for convenience of description. In actual applications, multiple steps can be combined or split, or executed in a different execution order, as long as the same technical effect can be achieved, all within the protection scope of the present invention. For example, steps S5 and S6 form a closed loop of iterative optimization and can be executed multiple times.

[0170] Embodiment 2 This embodiment provides a method for extracting merchant entity names, and its process may include the following steps, which are illustrated by specific examples: S1: Obtain the ringback tone recording data and convert the ringback tone recording data into a text content set.

[0171] In this step, the system first obtains a ringback tone recording from the outbound call platform. For example, an audio file containing the content "Hello, welcome to XX Cake Shop. We have new products on the market today..." is obtained. Then, the speech recognition (ASR) engine (such as Google Speech-to-Text) is called to process this audio file. The ASR engine analyzes the audio signal and converts it into initial text, which may be output as: "Hello, welcome to XX Cake Shop. We have new products on the market today".

[0172] Subsequently, basic text cleaning and normalization are performed. For example, punctuation is added and obvious recognition errors are corrected (if the ASR misidentifies "Cake Shop" as "Cake House", it may be corrected according to the language model at this stage, but more complex errors are left for subsequent processing). Finally, a set of normalized text content is obtained, which is the output of this example: "Hello, welcome to XX Cake Shop. We have new products on the market today...". This text set will be used as the input for the next step.

[0173] S2: Perform word segmentation and stop word removal on the set of text content to obtain a set of keywords, and label the set of keywords to obtain a labeled corpus.

[0174] Continuing with the text "Hello, welcome to XX Cake Shop. We have new products on the market today..." output from the previous step, in this step, a Chinese word segmentation tool (such as jieba) is first used for word segmentation to obtain a sequence of words: ["Hello", ",", "welcome to", "Chenxi", "cake", "shop", ",", "we", "today", "have", "new products", "on the market", "..."] (here it is assumed that "Cake Shop" is segmented into "cake" and "shop").

[0175] Then, according to a predefined stop word list (including "Hello", ",", "welcome to", "we", "today", "have", "...", etc.), these words that are not very meaningful for identifying the merchant name are removed. The remaining set of keywords is: ["Chenxi", "cake", "shop", "new products", "on the market"]. Then, the system may identify "XX Cake Shop" as a candidate name based on part of speech ("Chenxi", "cake", "shop" are noun components) and position.

[0176] Finally, in the manual annotation stage, the annotator confirms that "XX Cake Shop" is the correct merchant name on the annotation platform (such as doccano), and uses the B-NAME / I-NAME tagging method for annotation. The result is: Chenxi / B-NAME Cake / I-NAME Shop / I-NAME. Pair the original text with this tag sequence to form an annotated data: ("Hello, welcome to XX Cake Shop. We have new products on the market today...", ["O", "O", "O", "B-NAME", "I-NAME", "I-NAME", "O","O", "O", "O", "O", "O", "O"]). This data, together with a large number of other similar annotated data, jointly constitutes the annotated corpus for model training.

[0177] S3: Train a pre-trained large Chinese model based on the annotated corpus to obtain a merchant name extraction model.

[0178] In this step, use the annotated corpus constructed in the previous step, which contains a large number of annotated samples (such as examples containing "XX Cake Shop"), to train a named entity recognition (NER) model based on a pre-trained large Chinese model (such as "bert-base-chinese").

[0179] First, divide the corpus into a training set (e.g., 80%) and a test set (e.g., 20%). Then, input the training set data in batches (e.g., 32 samples per batch) into the configured BERT-NER model. The model predicts labels through forward propagation, calculates the loss between the predicted labels and the true labels (such as B-NAME, I-NAME, O) (e.g., using cross-entropy loss), and then updates the model parameters (weights and biases) according to the loss through the backpropagation algorithm. The optimizer (such as AdamW) and the learning rate (such as 3e-5) control the process of parameter update. This process is repeated for multiple rounds (epochs, e.g., 3 rounds).

[0180] After training is completed, use the test set to evaluate the performance of the model. For example, calculate the F1 score to be 95%. According to the preset strategy (such as selecting the model checkpoint with the highest F1 score on the validation set), finally output a trained model file that can recognize merchant names (such as merchant_ner_model.bin).

[0181] S4: Input the text to be processed into the merchant name extraction model to extract merchant name entities and obtain a merchant name extraction result set.

[0182] When the system receives a new, unlabeled ringback tone text, for example: "Thank you for calling Happy Pet Hospital. Please press 1 to transfer...". First, perform the same standardization process as during training on this text. Then, input the processed text into the merchant_ner_model.bin model trained in step S3. The model makes predictions for each character (or word) in the text and outputs a sequence of tags. For example, it might predict: Thank / O You / O Call / O Electric / O Happy / B-NAME Pet / I-NAME Hospital / I-NAME, / O Please / O Press / O 1 / O Transfer / O... / O.

[0183] The system then parses this sequence of tags and identifies the consecutive B-NAME and I-NAME tags starting from "Happy" and ending at "Hospital". Finally, based on this position information, extract the corresponding text fragment "Happy Pet Hospital" from the original text. Put all the extracted entities (in this case, only one) into a set to form the final merchant name extraction result set: {"Happy Pet Hospital"}.

[0184] S5: Verify the merchant name extraction result set and obtain optimization feedback information.

[0185] The system submits the result "Happy Pet Hospital" extracted in step S4 to the verification process. The verification can be manual review or comparison with a known merchant database. Assume the verification confirms that "Happy Pet Hospital" is correct. However, for another input "Please contact abcd Consulting Company for details", the model might wrongly extract only "abcd" and miss "Consulting Company". The verification process discovers this error and marks it as an abnormal sample (prediction: "abcd", actual: "abcd Consulting Company").

[0186] Meanwhile, the system will count the overall performance metrics over a period of time. For example, it is found that the recognition accuracy of merchant names containing suffixes such as "Company" and "Center" is relatively low. This information - including specific error samples and performance bottleneck analysis (such as the model's poor handling of long names or specific suffixes) - together constitutes the optimization feedback information. This information indicates the need to supplement training data or adjust the model specifically.

[0187] S6: According to the optimization feedback information, adjust the parameters of the merchant name extraction model and output the optimized merchant name extraction model for extracting merchant entity names.

[0188] After receiving the optimization feedback information in step S5, the development team or the automated system will take actions. For example, the sample with the recognized error "Please contact abcd consulting company for details" (annotated as "abcd / B-NAME cd / I-NAME consulting / I-NAME company / I-NAME") and more similar long names or samples containing specific suffixes will be added to the annotation corpus. Then, based on the performance bottleneck analysis, it is decided that the model structure may need to be adjusted (such as increasing the number of model layers) or the training strategy (such as adjusting the learning rate, increasing the number of training rounds). The model is retrained using the updated corpus and the adjusted strategy. Suppose the newly trained model version (e.g., v1.4) performs better on the validation set, especially being able to correctly identify samples such as "abcd consulting company", and the overall F1 score is improved from 95% to 96.5%. Then, this optimized model version v1.4 will be selected to replace the old version model currently in use online for subsequent merchant name extraction tasks. This "validation - feedback - optimization - deployment" loop will continue to continuously improve the model performance.

[0189] Embodiment 3 Refer to Figure 2 , the present invention also provides an extraction device 100 for merchant entity names, which can implement the method in the above embodiment. This device can be an independent hardware device or a software module running on a general computing device (such as a server, a computer). The device 100 may include: A data acquisition module 110, configured to acquire the ringback tone recording data and convert it into a text content set through speech recognition technology.

[0190] A data processing module 120, configured to preprocess the text content set, such as (optionally semantic-aware compression), word segmentation, stop word removal, to obtain a keyword set, and label the keyword set manually or semi-automatically (such as using B-NAME / I-NAME tags), and finally obtain an annotation corpus.

[0191] A model training module 130, configured to, based on the annotation corpus, (optionally perform data imbalance processing first), and then train the pre-trained Chinese large model (including data partitioning, model construction, parameter configuration, training, evaluation), (optionally perform inference and metric adjustment in combination with DCMN / POMDP), and finally obtain a merchant name extraction model.

[0192] A name extraction module 140, configured to receive the text to be processed, (optionally perform normalization processing), use the trained merchant name extraction model for prediction, identify the entity boundary, and extract the merchant name entity, and output a merchant name extraction result set.

[0193] A verification module 150 is used to verify the result set output by the name extraction module. For example, it evaluates using a verification data set, identifies abnormal samples, analyzes performance bottlenecks, and generates optimization feedback information including performance metrics and error cases.

[0194] An iterative optimization module 160 is used to adjust the parameters of the merchant name extraction model according to the optimization feedback information generated by the verification module (possibly by triggering the model training module for retraining), outputs a merchant name extraction model with better performance, and provides it for the name extraction module to use.

[0195] It should be understood that Figure 2 The shown device structure is only an example. The functions of each module can be combined or split, and can also be implemented in ways of software, hardware, or a combination of software and hardware. For example, the functions of the model training module and the iterative optimization module can be completed in an offline environment, while the data acquisition, data processing (part), name extraction, and verification (part) modules can run online in real time. These modules can communicate and transfer data through internal interfaces or buses.

[0196] In summary, the present invention significantly improves the accuracy, robustness, and efficiency of extracting merchant entity names from ringback tone recording data by combining a pre-trained large model, refined annotation, targeted data processing (such as semantic compression, imbalance processing), uncertainty modeling, and a closed-loop iterative optimization mechanism, and has high practical value.

[0197] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

[0198] The embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of a method for extracting a merchant entity name in the above method embodiments. Among them, the storage medium can be a volatile or non-volatile computer-readable storage medium.

[0199] In addition, the embodiments of the present disclosure also provide a computer program product, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of a method for extracting a merchant entity name provided in any of the above embodiments of the present disclosure. For details, refer to the above method embodiments and will not be elaborated here.

[0200] Among them, the above computer program product can be specifically implemented in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium, which can be a volatile or non-volatile computer-readable storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0201] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and apparatuses can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. In several embodiments provided by the present disclosure, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection can be through some communication interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0202] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0203] In addition, in each embodiment of the present disclosure, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0204] When the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0205] Finally, it should be noted that the above-mentioned embodiments are only specific implementation manners of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A method for extracting the merchant entity name, characterized in that, Including: Obtain the ringback tone recording data and convert the ringback tone recording data into a text content set; Perform word segmentation processing and stop word removal processing on the text content set to obtain a keyword set, and perform annotation on the keyword set to obtain an annotated corpus; Train a pre-trained large Chinese model based on the annotated corpus to obtain a merchant name extraction model; Input the text to be processed into the merchant name extraction model, extract the merchant name entity, and obtain a merchant name extraction result set; Verify the merchant name extraction result set to obtain optimization feedback information; Adjust the parameters of the merchant name extraction model according to the optimization feedback information, and output an optimized merchant name extraction model for extracting merchant entity names.

2. The extraction method of the merchant entity name according to claim 1, characterized in that The performing word segmentation processing and stop word removal processing on the text content set includes: Use a word segmentation tool to process the text content set to obtain a word sequence; Remove stop words from the word sequence to form the keyword set; Use the keyword set to extract entity information to obtain a candidate name set; Use B-NAME and I-NAME tags to process the candidate name set to construct the annotated corpus.

3. The extraction method of the merchant entity name according to claim 1, characterized in that, The training the pre-trained large Chinese model based on the annotated corpus includes: Use a data partitioning algorithm to process the annotated corpus to obtain a training set and a test set; Train the pre-trained large Chinese model with the training set and output an initial model; Use the test set to evaluate the performance of the initial model to obtain evaluation metrics; Adjust the parameters of the initial model according to the evaluation metrics to obtain the merchant name extraction model.

4. The extraction method of the merchant entity name according to claim 3, wherein Before adjusting the parameters of the initial model according to the evaluation metrics, it further includes: Adjusting the evaluation metrics includes: using the initial model to generate a label prediction result as output data; Construct a hybrid model based on a deep belief Markov network and POMDP; Input the output data into the hybrid model to obtain a probability distribution sequence; Perform state inference calculation on the probability distribution sequence through the hybrid model to obtain an adjusted evaluation metric; Use the adjusted evaluation metric to adjust the parameters of the initial model.

5. The extraction method of the merchant entity name according to claim 1, characterized in that, Before performing word segmentation processing on the text content set, it includes: Use a time series encoder to process the text content set to obtain a time series representation; Perform data dimensionality reduction processing on the time series representation and output a compressed sequence; Construct a semantic analysis model and use the semantic analysis model to process the compressed sequence to obtain semantic features related to the merchant name; Among them, the performing word segmentation processing on the text content set includes: using the semantic features related to the merchant name to guide the word segmentation processing.

6. The extraction method of the merchant entity name according to claim 5, wherein, The performing data dimensionality reduction processing on the time series representation includes: Construct a semantic importance calculation model and use the semantic importance calculation model to process the time series representation to obtain an importance distribution; According to the importance distribution, calculate the differential sampling frequency of the time series representation, and output a target sampling parameter including a sampling rate and a sampling point position. Downsample the time series representation using the target sampling parameters to obtain a sampled sequence; Divide the sampled sequence into multiple semantic units, and perform semantic integrity detection on the multiple semantic units, and output the compressed sequence.

7. The extraction method of the merchant entity name according to claim 1, characterized in that Before training the pre-trained Chinese large model based on the annotated corpus, it includes: Perform data distribution calculation on the annotated corpus to obtain the frequency characteristics of merchant name entity categories; Use the frequency characteristics of the merchant name entity category to calculate the data imbalance degree in the annotated corpus; Generate sample weight coefficients according to the data imbalance degree; Use the sample weight coefficients to construct the training loss function of the pre-trained Chinese large model, and determine the parameter settings for model training.

8. The method for extracting the merchant entity name according to claim 7, wherein, The calculating the data imbalance degree in the annotated corpus by using the frequency characteristics of the merchant name entity category includes: Classify the frequency characteristics of the merchant name entity category according to the distribution characteristics to form multiple evaluation units; Obtain the feature subset of the merchant name entity category from the multiple evaluation units; Calculate the gradient change amount of the feature subset and output the gradient optimization direction; Set the training loss function parameters according to the gradient optimization direction and the data imbalance degree to generate a training objective function; Use the training objective function to perform parameter adaptive adjustment on the pre-trained Chinese large model to generate sample weight coefficients inversely proportional to the data imbalance degree.

9. The extraction method of the merchant entity name according to claim 1, wherein The inputting the text to be processed into the merchant name extraction model includes: Perform text normalization processing on the text to be processed to obtain a normalized text; Input the normalized text into the merchant name extraction model to obtain a character probability distribution; Use the character probability distribution to identify the entity boundary positions and output a position marking sequence; Extract the corresponding text segments from the normalized text according to the position marking sequence to generate the merchant name extraction result set.

10. The extraction method of the merchant entity name according to claim 1, characterized in that, The verifying the merchant name extraction result set includes: Evaluate the merchant name extraction result set using a validation data set to obtain an abnormal sample set; Expand the annotated corpus according to the abnormal sample set to form an enhanced training set; Analyze the model performance bottleneck based on the enhanced training set, determine the model structure parameters to be optimized, and generate a model tuning scheme; Execute the model tuning scheme to train the merchant name extraction model, and calculate the recognition rate and accuracy rate of the model on the validation data set to form optimization feedback information including performance indicators and error cases.

11. An extraction device for the merchant entity name, characterized in that, It includes: A data acquisition module for acquiring ringback tone recording data and converting the ringback tone recording data into a text content set; A data processing module for performing word segmentation processing and stop word removal processing on the text content set to obtain a keyword set, and performing annotation based on the keyword set to obtain an annotated corpus; A model training module for training a pre-trained Chinese large model based on the annotated corpus to obtain a merchant name extraction model; A name extraction module for inputting the text to be processed into the merchant name extraction model, extracting merchant name entities, and obtaining a merchant name extraction result set; A verification module, which is used to verify the merchant name extraction result set and obtain optimization feedback information; An iterative optimization module, which is used to adjust the parameters of the merchant name extraction model according to the optimization feedback information and output an optimized merchant name extraction model for extracting merchant entity names.