Cultural value question and answer model training method and device, equipment and medium

By constructing a cultural value question-answering model through vectorization and clustering algorithms, and combining a reward mechanism and a group relative optimization strategy, the problem of insufficient consistency in cross-cultural question-answering models is solved, and better question-answering performance and generation consistency are achieved under different cultures.

CN121808392APending Publication Date: 2026-04-07CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies in cross-cultural question answering models suffer from insufficient cultural consistency, easy mismatch in fine-grained and low-resource cultural scenarios, and weak cross-cultural generalization ability, especially in open generation or out-of-distribution scenarios where alignment drift is prone to occur.

Method used

By obtaining the vectorized processing of the training questions, the target cluster and cluster centers are generated using a clustering algorithm, a hierarchical statement list is constructed, and the question-answering model is updated to improve consistency by combining a reward mechanism and a group relative optimization strategy model.

Benefits of technology

It improves the performance of question-answering models across different cultures, achieves more stable cross-cultural alignment and generation consistency, reduces data costs, and enhances the applicability to low-resource cultures and the performance of generation tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121808392A_ABST
    Figure CN121808392A_ABST
Patent Text Reader

Abstract

The invention provides a culture value question and answer model training method and device, equipment and a medium, and the method comprises the steps: obtaining a training question, determining a training cluster corresponding to the training question according to a question vector corresponding to the training question and a cluster center of each target cluster, and obtaining a culture value question and answer model according to the training cluster and a training country corresponding to the training question. Determining a training statement set from the hierarchical statement list; according to the training statement set, a target question and answer model, a preset question and answer frequency, the training question and a corresponding training country, determining each candidate output and a corresponding reward value; and according to each reward value, determining a relative reward advantage, and through a group relative optimization strategy model, according to the training question, each candidate output and the relative reward advantage, updating the target question and answer model. Through the technical scheme of the invention, the model effect of the target question and answer model is improved, so that the target question and answer model has better performance under different cultures.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to a cultural value question and answer model training method and device, equipment and medium. BACKGROUND

[0002] At present, the self-alignment method based on prompt engineering can guide the model to generate answers conforming to the culture in the reply stage by retrieving and implanting target cultural value statements and a small amount of examples. The technical route of semantic expansion and data enhancement based on world value survey data and subsequent supervised fine-tuning can realize question and answer generation and alignment for multiple national cultures.

[0003] However, the method based on the existing prompt engineering is highly sensitive to the instruction understanding and context retrieval of the model, and it is difficult to maintain stable consistency in fine-grained or unfamiliar cultural contexts. The reason is that its alignment mechanism mainly relies on temporary context constraints rather than learnable long-term parameterized representations. The supervised fine-tuning method based on data enhancement relies significantly on large-scale and balanced cross-cultural annotation samples. Low-resource cultures are prone to data sparsity and class bias, resulting in overfitting and insufficient cross-cultural generalization ability. The reason is that the training target is mainly based on imitative classification / generation loss, and lacks cross-cultural invariance constraints. These two methods generally lack explicit modeling of "cultural norms" and interpretable internalization processes, and more are fitting of example distribution and discourse patterns, leading to alignment drift problems in open generation or out-of-distribution scenarios. SUMMARY

[0004] In view of the above-mentioned defects or deficiencies in the prior art, the present application aims to provide a cultural value question and answer model training method, device, equipment and medium to solve the problems of insufficient cultural consistency, easy mismatch in fine-grained and low-resource cultural scenarios, weak cross-cultural generalization ability, and difficulty in maintaining the model scale in multiple cultural values.

[0005] The cultural value question and answer model training method provided by the present application comprises the following steps: Obtaining a training question, determining a training cluster corresponding to the training question according to a question vector corresponding to the training question and cluster centers of each target cluster, and determining a training statement set from a hierarchical statement list according to the training cluster and a training country corresponding to the training question; wherein the target cluster and the cluster center are clustering results of question vectors corresponding to each survey question; and the hierarchical statement list is a cultural value statement list constructed by taking each country and each target cluster as an index; Determining each candidate output and a corresponding reward value, respectively, according to the training statement set, a target question and answer model, a preset question and answer number, the training question, and the corresponding training country; According to each reward value, a relative reward advantage is determined, and a group relative optimization strategy model is used to update the target question and answer model according to the training question, each candidate output, and the relative reward advantage.

[0006] According to the technical scheme provided in the embodiments of the present application, optionally, each target cluster, the cluster center of each target cluster, and the hierarchical statement list are determined based on the following manner: According to the question vector corresponding to each survey question in the cultural value question and answer set, a clustering operation is performed to obtain each target cluster and the corresponding cluster center; The cultural value question and answer set is processed with each target cluster and each country in the cultural value question and answer set as an index to construct a hierarchical statement list. The cultural value question and answer set includes the basic information of each survey question, the answering frequency of each country corresponding to each survey question, the answer selection proportion of each country on the survey question, the mainstream answer, and the mainstream intensity.

[0007] According to the technical scheme provided in the embodiments of the present application, optionally, before the clustering operation is performed according to the question vector corresponding to each survey question in the cultural value question and answer set to obtain each target cluster and the corresponding cluster center, the following steps are further included: For each survey question in the preselected survey questionnaire, the answering frequency of each country corresponding to the survey question and the answer selection proportion of each country on the survey question are determined according to the survey answers of each country to the survey question; For each two to-be-calculated countries, the answer similarity of the two to-be-calculated countries on the survey question is determined according to the answer selection proportion of the two to-be-calculated countries on the survey question; The similarity threshold of the two to-be-calculated countries on the survey question is determined according to the answering frequency of the two to-be-calculated countries corresponding to the survey question; In response to the answer similarity being greater than the similarity threshold, the survey answers corresponding to the two to-be-calculated countries on the survey question are filtered; The answering frequency of each country corresponding to the survey question is determined after filtering, and the question and answer standardized data of the survey question is constructed according to the basic information of the survey question, the filtered countries, the answering frequency of the filtered countries corresponding to the survey question, the mainstream answer corresponding to the maximum value of the answer selection proportion, and the mainstream intensity. The cultural value question and answer set is constructed according to the question and answer standardized data of each survey question in the preselected survey questionnaire.

[0008] According to the technical scheme provided in the embodiments of the present application, optionally, the following steps are further included: For each survey question in the pre-selected questionnaire, the cultural value theme and theme confidence level corresponding to the survey question are determined based on a preset large language model. If the confidence level of the topic is lower than the preset confidence level, the survey question is deleted to update the pre-selected survey questionnaire.

[0009] According to the technical solution provided in the embodiments of this application, optionally, determining each candidate output and its corresponding reward value based on the training statement set, the target question-answering model, the preset number of question-answering attempts, the training question, and the corresponding training country includes: According to the preset number of questions and answers, the training questions and the corresponding training countries are input into the target question-answering model to obtain each candidate output; For each candidate output, a consensus result is determined based on the candidate output and the training statement set, and the reward value corresponding to the candidate output is determined based on the consensus result.

[0010] According to the technical solution provided in the embodiments of this application, optionally, the step of determining each consistency result based on the candidate output and the training statement set, and determining the reward value corresponding to the candidate output based on each consistency result, includes: For each target question in the training statement set, the target value statement and the candidate output are evaluated for consistency based on a pre-defined large language model to determine the consistency result corresponding to the target question; wherein, the target value statement corresponding to the target question includes the frequency of responses, mainstream answers, and mainstream intensity of responses from the training countries to the target question; Determine whether at least one of the consistent results is contradictory; If at least one consistent result is contradictory, then the reward value corresponding to the candidate output is determined to be the first value; If all consistent results are not contradictory, then determine whether there is at least one supporting result among the consistent results; If at least one consistent result is supported, then the reward value corresponding to the candidate output is determined to be the second value; If none of the consistency results support the candidate output, then the reward value corresponding to the candidate output is determined to be the third value. Wherein, the first value is less than the third value, and the third value is less than the second value.

[0011] According to the technical solution provided in the embodiments of this application, optionally, determining the training cluster corresponding to the training problem based on the problem vector corresponding to the training problem and the cluster center of each target cluster includes: For each target cluster, the cluster distance is determined based on the question vector corresponding to the training question and the cluster center of the target cluster, and the minimum value among the cluster distances is taken as the minimum distance. If the minimum distance is less than or equal to a preset distance, then the target cluster corresponding to the minimum distance is taken as the training cluster corresponding to the training problem; If the minimum distance is greater than the preset distance, then the training problem is determined to be an irrelevant problem.

[0012] This application embodiment also provides a cultural value question-answering model training device, the device comprising: The information extraction module is used to obtain training questions, determine the training clusters corresponding to the training questions based on the question vectors corresponding to the training questions and the cluster centers of each target cluster, and determine the training statement set from the hierarchical statement list based on the training clusters and the training countries corresponding to the training questions; wherein, the target clusters and cluster centers are the clustering results of the question vectors corresponding to each survey question; the hierarchical statement list is a list of cultural value statements constructed with each country and each target cluster as indexes; The reward value calculation module is used to determine each candidate output and its corresponding reward value based on the training statement set, the target question-answering model, the preset number of question-answering attempts, the training question, and the corresponding training country. The model training and update module is used to determine the relative reward advantage based on each reward value, and update the target question answering model based on the training question, each candidate output, and the relative reward advantage through a group relative optimization strategy model.

[0013] This application also provides an electronic device, the electronic device comprising: Processor and memory; The processor executes the steps of the cultural value question-answering model training method as described in any embodiment by calling the program or instructions stored in the memory.

[0014] This application also provides a computer-readable storage medium storing a program or instructions that cause a computer to perform the steps of the cultural value question-answering model training method as described in any embodiment.

[0015] In summary, this application proposes a training method for a cultural value question-answering model. By acquiring training questions, determining the training clusters corresponding to the training questions based on the question vectors and cluster centers of each target cluster, and then determining the training statement set from a hierarchical statement list based on the training clusters and the corresponding training countries, the method retrieves surveyed information related to the training questions and countries. Furthermore, based on the training statement set, the target question-answering model, the preset number of question-answers, the training questions, and the corresponding training countries, the method determines each candidate output and its corresponding reward value. This is combined with the surveyed information to analyze the accuracy of each candidate output from the target question-answering model and provide corresponding reward values. Finally, based on each reward value, the method determines the relative reward advantage and updates the target question-answering model using a group relative optimization strategy model, based on the training questions, candidate outputs, and relative reward advantage. This improves the model's performance, enabling it to achieve better results across different cultures. Attached Figure Description

[0016] Figure 1 This is a flowchart of a cultural value question-answering model training method provided in an embodiment of this application; Figure 2 This is a flowchart of another cultural value question-answering model training method provided in the embodiments of this application; Figure 3 This is a schematic diagram illustrating the performance of the cultural reinforcement learning method and supervised fine-tuning method provided in the embodiments of this application on open-ended problems; Figure 4 This is a schematic diagram of the structure of a cultural value question-and-answer model training device provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0017] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.

[0018] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0019] Figure 1 This is a flowchart illustrating a method for training a cultural value question-and-answer model according to an embodiment of this application. See also... Figure 1 The training method for this cultural value question-and-answer model specifically includes: S110. Obtain the training problem. Based on the problem vector corresponding to the training problem and the cluster center of each target cluster, determine the training cluster corresponding to the training problem. Based on the training cluster and the training country corresponding to the training problem, determine the training statement set from the hierarchical statement list.

[0020] The training question is a culture-related question used to train the target question-answering model. It does not require collecting real answers for this training question; the model is designed for cultural value question answering. The question vector is the vector obtained after vectorizing the training question. The target cluster and cluster center are the clustering results of the question vectors corresponding to each survey question. The target cluster is the cluster obtained by clustering the vectorized survey questions in the hierarchical statement list. The cluster center is the center vector of each target cluster. The training cluster is the target cluster closest to the question vector. The training country is the country specifically trained for the training question during the model training process. It can be understood that the training process in this example has generalization potential, and the trained model can improve performance for other countries as well. The hierarchical statement list is a list of cultural value statements constructed using each country and target cluster as indexes. The cultural value statement list is a statistical list obtained by integrating the question stems, answer codes, and selection results for each country and each survey question. The training statement set is a partial set of information retrieved from the hierarchical statement list according to the training cluster and training country.

[0021] Specifically, the system receives training questions input by the user for further training of the target question-answering model. These training questions are vectorized to obtain corresponding question vectors. The distance between these question vectors and the cluster centers of each target cluster is calculated, and the target cluster with the closest distance is identified as the training cluster corresponding to the training question. The system also receives the training countries corresponding to the training questions. Using the training clusters and training countries as indices, the system retrieves the corresponding information from the hierarchical statement list, forming a training statement set.

[0022] Understandably, during the generation of target clusters, each survey question undergoes vectorization. To improve clustering performance, abstraction is performed before vectorization; therefore, the training questions also undergo the same abstraction process. Taking survey questions as an example, the abstraction process specifically involves rewriting the question stems into standardized "cultural concept" phrases, including removing stop words, removing question format structures, and normalizing entity names. The vectorization process involves encoding the questions into a pre-defined vector space using a unified embedding model, essentially establishing a question-vector mapping.

[0023] Based on the above example, each target cluster, the cluster center of each target cluster, and the hierarchical statement list are determined in the following manner: Based on the question vector corresponding to each survey question in the cultural value question and answer set, a clustering operation is performed to obtain each target cluster and its corresponding cluster center; Using each target cluster and each country in the cultural value question and answer set as an index, the cultural value question and answer set is processed to construct a hierarchical statement list.

[0024] The cultural value question and answer set includes basic information for each survey question, as well as the corresponding countries, response frequency, dominant answers, and dominant strength for each country. Basic information includes the question number, answer code, and answer content. Response frequency refers to the percentage of responses from different countries. Dominant answers are the answer codes and content with the highest selection rate for a given country. The answer selection rate is the percentage of selections for different answer codes within each country. Dominant strength is the percentage of selections for the corresponding dominant answer.

[0025] Specifically, each survey question in the cultural value question-and-answer set is vectorized to obtain a corresponding question vector. Clustering these question vectors yields target clusters and their corresponding cluster centers. Using each target cluster and each country in the cultural value question-and-answer set as indexes, the set is hierarchically processed to create a hierarchical statement list—that is, a hierarchical index of {country→cluster→cultural value statement list}—supporting fast retrieval by country and cluster.

[0026] Understandably, the clustering process can use the K-means clustering algorithm. For the number of clusters K, iterate through the clustering results of K, draw the inter-cluster distance-K plot and the silhouette coefficient-K plot, and find the inflection point by comprehensively judging the two plots to obtain the optimal K value for clustering.

[0027] Building upon the example above, before performing clustering operations based on the question vector corresponding to each survey question in the cultural value question-and-answer set to obtain each target cluster and its corresponding cluster center, there is also a process for constructing the cultural value question-and-answer set. This can be done in the following way: For each survey question in the pre-selected questionnaire, based on the survey answers from each country, determine the corresponding response frequency and the proportion of answers selected by each country for each survey question; For each pair of countries to be counted, the similarity of their answers to the survey questions is determined based on the proportion of their responses to the survey questions. Based on the response frequencies of the two countries to be calculated on the survey questions, a similarity threshold for the two countries to be calculated on the survey questions is determined. If the similarity of the answers is greater than the similarity threshold, the corresponding survey answers for the two countries to be calculated in the survey question will be filtered. Determine the response frequency of each country on the survey questions after filtering. Based on the basic information of the survey questions, the filtered countries, the response frequency of each filtered country, the mainstream answer corresponding to the maximum value in the answer selection ratio, and the mainstream strength, construct standardized data of the survey questions and responses. Based on the standardized data of the questions and answers for each survey question in the pre-selected questionnaire, a set of questions and answers on cultural values ​​was constructed.

[0028] The pre-selection questionnaire can be an open-source questionnaire investigating cultural values, such as data from a global values ​​survey, including survey questions and answers from various countries, and the results of the responses regarding the selection of answer options and content. Alternatively, the pre-selection questionnaire can be other cultural value-related questionnaires and answers conducted by country. Each pair of countries to be counted consists of two distinct countries that responded to the pre-selection questionnaire. Response similarity is the degree of similarity in the proportion of answer selections for the survey questions between the two countries to be counted. The similarity threshold is a numerical value used to measure whether the responses from two countries to be counted are excessively similar. Standardized question-and-answer data is the data obtained by standardizing the filtered survey questions and related information.

[0029] Specifically, for each survey question in the pre-selected questionnaire, the survey answers for each country are statistically analyzed to determine the number of answers for each country, the percentage of answers, and the corresponding frequency of responses for each country. Furthermore, the proportion of answer selections for each country on that question is statistically calculated. For every two countries to be analyzed, the proportion of answer selections for the survey question is obtained, and similarity is calculated based on these proportions, for example, using Jensen-Shannon divergence (JS divergence), to obtain the similarity of answers between the two countries on that survey question. The sum of the response frequencies for the two countries on that survey question is obtained, and a similarity threshold is determined based on this sum of frequencies, for example, using a pre-constructed list or a pre-constructed functional relationship. This threshold is then used as the similarity threshold between the two countries on that survey question. If the answer similarity exceeds a similarity threshold, it indicates that the two countries being calculated are too similar on this survey question and lack cross-national diversity. Therefore, the survey answers for these two countries on this survey question are filtered and deleted. After filtering using similarity, the frequency of responses for each filtered country on the survey question is updated. Furthermore, the basic information of the survey question, the filtered countries, the frequency of responses for each filtered country, the dominant answer corresponding to the maximum answer selection ratio, and the dominant strength are standardized according to a preset field standardization structure to construct standardized question-and-answer data for the survey question. The standardized question-and-answer data for each survey question in the pre-selected questionnaire are combined to construct a cultural value question-and-answer set.

[0030] Building upon the above example, cultural value-related screening can be performed on the survey questions in the pre-selected questionnaire. Specifically, this could include: For each survey question in the pre-selected questionnaire, the cultural value theme and theme confidence level corresponding to the survey question are determined based on the pre-set large language model. If the topic confidence level is lower than the preset confidence level, the survey question will be deleted and the pre-selected questionnaire will be updated.

[0031] The pre-selected large language model, such as GPT-4O, is used to evaluate cultural value topics. Cultural value topics can be various themes related to cultural values, such as morality / work. Topic confidence is the credibility of the cultural value topics determined through analysis of the survey questions. Pre-set reliability is a pre-determined numerical value used to measure whether the survey questions are sufficiently relevant to cultural values.

[0032] Specifically, for each survey question in the pre-selected questionnaire, the question is input into a pre-defined large language model for analysis of cultural value themes. This model outputs the corresponding cultural value theme and the theme confidence level. If the theme confidence level is lower than the pre-set confidence level, it indicates that the credibility of the cultural value theme most relevant to the survey question is also low. This confirms that the survey question has a poor correlation with cultural values, and the survey question is deleted to update the pre-selected questionnaire.

[0033] S120. Based on the training statement set, the target question-answering model, the preset number of question-answering attempts, the training questions, and the corresponding training countries, determine each candidate output and its corresponding reward value.

[0034] The preset number of question-answering attempts is a pre-defined number of times the target question-answering model will answer a training question and its corresponding training country, such as 8 times. The candidate output is the answer obtained each time the target question-answering model is used to answer the training question and its corresponding training country. The reward value describes whether the candidate output matches the corresponding set of training statements.

[0035] Specifically, the training questions and corresponding training countries are input into the target question-answering model according to a preset number of question-answering attempts, and a candidate answer is obtained each time. For each candidate answer, it is matched and evaluated with the determined set of training statements, and the reward value corresponding to the candidate answer is determined according to the matching result.

[0036] S130. Based on each reward value, determine the relative reward advantage, and update the target question-answering model using a group relative optimization strategy model, based on the training question, each candidate output, and the relative reward advantage.

[0037] Among them, the group relative optimization strategy model is a reinforcement learning algorithm specifically designed to enhance the reasoning ability in large language models. It optimizes the model by evaluating mutually related response groups (each candidate output). The relative reward advantage is the advantage of each candidate answer relative to the group of candidate answers.

[0038] Specifically, a reward vector is constructed based on each reward value. Using existing methods for calculating relative reward advantage, the relative reward advantage corresponding to each candidate answer can be obtained. Furthermore, through a group relative optimization strategy model, the target question answering model is optimized using the training question, each candidate output, and the relative reward advantage to update and obtain a better target question answering model.

[0039] The group relative optimization policy model is optimized by maximizing the following objective, which includes both a reward-driven policy update term and a KL divergence regularization term, to maintain training stability and avoid deviating too far from the reference model:

[0040]

[0041] in, It is the objective function, and E is the expectation. It follows a certain probability distribution. q represents the model parameters that need to be optimized in the target question-answering model, and q is the training question. It is the probability distribution corresponding to the set of cultural value questions and answers, used to describe the probability of different questions appearing. is the i-th candidate answer output by the target question-answering model before optimization, G is the preset number of question-answering attempts, and min is the minimum value. This is the original question-answering model. It represents the probability distribution of each answer given a training question q by the target question-answering model before optimization. This is the optimized target question-answering model, where t is the time step. Each step before time step t , It represents the relative reward advantage at time step t among the i-th candidate outputs. `clip` is a classic pruning function in Proximal Policy Optimization (PPO). This is the cutting factor. It is the KL divergence between the target question-answering models before and after optimization. These are preset hyperparameters. It is a reference model, usually the target question-answering model before optimization.

[0042] Optionally, the updated target question-answering model can be used to predict questions and answers related to cultural values. Only the country and the question need to be input. If only the question is input, the answer situation of each country can be predicted.

[0043] Optionally, the above example supports two training paradigms: One-for-One and One-for-All. One-for-One means that each country's culture trains an independent model, which only performs normative alignment on the target culture of a single country. One-for-All means that multiple countries' culture normative constraints are learned simultaneously in the same policy model, and unified alignment of multiple cultures is achieved by switching the target country or culture label.

[0044] The cultural value question-answering model training method provided in this application obtains training questions, determines training clusters corresponding to the training questions based on the question vectors corresponding to the training questions and the cluster centers of each target cluster, determines a set of training statements from a hierarchical statement list based on the training clusters and the training countries corresponding to the training questions, and retrieves surveyed information related to the training questions and training countries. Then, based on the set of training statements, the target question-answering model, the preset number of question-answering attempts, the training questions, and the corresponding training countries, each candidate output and its corresponding reward value are determined. The accuracy of each candidate output of the target question-answering model is analyzed in conjunction with the surveyed information, and corresponding reward values ​​are given. Finally, based on each reward value, the relative reward advantage is determined, and the target question-answering model is updated using a group relative optimization strategy model based on the training questions, each candidate output, and the relative reward advantage. This improves the model performance of the target question-answering model, enabling it to perform better under different cultures.

[0045] Figure 2 This is a flowchart of another cultural value question-answering model training method provided in this application. Based on the above embodiments, the process of determining the training cluster corresponding to the training question and the process of determining each candidate output and its corresponding reward value are illustrated. See also... Figure 2 The training method for this cultural value question-and-answer model specifically includes: S210. Obtain the training problem. For each target cluster, determine the cluster distance based on the problem vector corresponding to the training problem and the cluster center of the target cluster, and take the minimum value among the cluster distances as the minimum distance.

[0046] Here, the cluster distance is the distance between the question vector corresponding to the training question and the cluster center of the target cluster. The minimum distance is the minimum value among all the cluster distances corresponding to the training question.

[0047] Specifically, the training problem is obtained and vectorized to obtain its corresponding problem vector. For each target cluster, the distance between the training problem's corresponding problem vector and the cluster center of the target cluster is calculated to obtain the problem cluster distance between the training problem and that target cluster. The minimum value among the problem cluster distances corresponding to the training problem across all target clusters is taken as the minimum distance.

[0048] S220. If the minimum distance is less than or equal to the preset distance, the target cluster corresponding to the minimum distance is taken as the training cluster corresponding to the training problem; if the minimum distance is greater than the preset distance, the training problem is determined to be an irrelevant problem.

[0049] The preset distance is used to measure whether a problem vector can be assigned to a corresponding target cluster. Irrelevant problems are training problems that are unrelated to each target cluster and are not used for subsequent training.

[0050] Specifically, if the minimum distance is less than or equal to the preset distance, it means that the problem vector of the training problem is close enough to the target cluster corresponding to the minimum distance, and can be assigned to that target cluster. Therefore, the target cluster corresponding to the minimum distance is taken as the training cluster corresponding to the training problem. If the minimum distance is greater than the preset distance, it means that the correlation between the training problem and each target cluster is relatively poor. Therefore, the training problem is determined to be an irrelevant problem.

[0051] S230. Based on the training clusters and the training countries corresponding to the training questions, determine the training statement set from the hierarchical statement list.

[0052] S240. According to the preset number of question-and-answer sessions, input the training questions and the corresponding training countries into the target question-and-answer model to obtain each candidate output.

[0053] S250. For each candidate output, determine the consistency results based on the candidate output and the training statement set, and determine the reward value corresponding to the candidate output based on the consistency results.

[0054] The consistency result is used to describe whether the candidate output matches each piece of information in the training statement set, and can include contradictions, support, and irrelevance.

[0055] Specifically, a pre-selected large language model is used as the external reward model for evaluation, such as GPT-4o-mini. For each candidate output, the candidate output and the training statement set are input into the external reward model. The model matches and evaluates the candidate output with each piece of information in the training statement set to obtain the consistency result between the candidate output and each piece of information. A comprehensive evaluation of the consistency results yields the reward value corresponding to the candidate output.

[0056] Based on the above example, the consistency results can be determined according to the candidate outputs and the training statement set in the following way: For each target question in the training statement set, the target value statement is evaluated for consistency with the candidate output based on the pre-defined large language model, and the consistency result corresponding to the target question is determined.

[0057] The target question is any of the survey questions included in the training statement set. The target value statement corresponding to the target question is the information in the training statement set that corresponds to the target question, including the frequency of responses, mainstream answers, and mainstream strength of the training countries to the target question.

[0058] Specifically, for each target value statement corresponding to the target question in the training statement set, a pre-set large language model is used to determine the consistency between the target value statement and the candidate output. This determines whether the target question and the target value statement match, which is the consistency result.

[0059] Based on the above example, the reward value corresponding to each candidate output can be determined according to the consistency results in the following way: Determine whether at least one of the consistent results is contradictory; If at least one consistent result is contradictory, then the reward value corresponding to the candidate output is determined to be the first value; If all consistent results are not contradictory, then determine whether there is at least one supporting result among the consistent results; If at least one consistent result supports the candidate output, then the reward value corresponding to the candidate output is determined to be the second value. If none of the consistency results support the candidate output, then the reward value corresponding to the candidate output is determined to be the third value.

[0060] In this system, the first value is less than the third value, and the third value is less than the second value. The first, second, and third values ​​are usually pre-set values. The first value is typically negative and used for punishment; the third value is typically zero and indicates no relevance or punishment; and the second value is typically positive and used for reward. For example, the first value could be -1, the second value could be 1, and the third value could be 0.

[0061] Specifically, the reward value is determined using a priority-based short-circuit ranking decision. This strategy possesses anti-dilution property (no contradiction can be diluted by majority support), monotonicity (highest sensitivity to contradiction determination), and short-circuit computation. Specifically, it checks whether at least one of the consensus results is contradictory. If at least one consensus result is contradictory, the candidate output is directly determined to be contradictory, and its corresponding reward value is the first value, such as -1. If none of the consensus results are contradictory, it further checks whether at least one of the consensus results is supportive. If at least one consensus result is supportive, the candidate output is directly determined to be supportive, and its corresponding reward value is the second value, such as 1. If none of the consensus results are supportive, it means that all consensus results are irrelevant, and the reward value corresponding to the candidate output is determined to be the third value, such as 0.

[0062] For example, the reward value corresponding to the candidate output can be calculated using the following formula:

[0063] in, Let G be the reward value corresponding to the i-th candidate output, where i ∈ [1, G], and G is the preset number of questions and answers. This represents the consistency result between the candidate output and the j-th target value statement in the training statement set.

[0064] S260. Based on each reward value, determine the relative reward advantage, and update the target question-answering model using a group relative optimization strategy model, based on the training question, each candidate output, and the relative reward advantage.

[0065] The technical solution described above can be embedded in the intelligent large language model of the vehicle host, so that users can ask and answer questions about cultural values ​​through the vehicle host and obtain reasonable answers.

[0066] The technical solution in this example can be summarized as follows: it includes the following two stages: Phase 1: Construction of the Value Norms Pool (A Collection of Cultural Value Questions and Answers) Based on the "World Values ​​Survey" as a pre-selected questionnaire, a value norm database corresponding to different cultures is constructed. The input survey questions are selected from the pre-selected questionnaire (e.g., "Do you agree that work should always take precedence, even if it means less free time?"). Questions that meet the criteria are that they must be relevant to cultural values ​​and that different countries have different responses to them. Then, from the selected questions, corresponding cultural value concepts are extracted (e.g., "Work takes precedence over leisure"). Finally, these are converted into value norms. This can be done by combining responses from different countries (e.g., "Strongly agree" in country X, "Disagree" in country Y, and "Agree" in country Z), converting "question + corresponding country response" into a value norm pool (standardized question-and-answer data) to clarify the value orientations of different cultures (e.g., "This culture values ​​work over leisure" corresponds to country X / Z; "This culture prioritizes personal time over work" corresponds to country Y).

[0067] Phase 2: Reward Mechanism Based on Normal Clustering Using a value norm pool, the model (target question-answering model) is optimized for group-based relative strategies. Each survey question is clustered to obtain target clusters and their corresponding cluster centers. Then, cultural value norms are matched. For the target participants (e.g., "K country participants"), the "K country value norms" (containing corresponding cultural value tendencies) in the value norm pool are retrieved and hierarchically organized according to each target cluster to obtain a hierarchical statement list. Training questions (such as the aforementioned work priority questions) are input into the model, and the model outputs candidate answers (e.g., "Yes, I agree"). Rewards are determined according to rules: if a candidate answer contradicts the corresponding cultural norm (training statement set), the reward value is -1; if the candidate answer supports the corresponding cultural norm, the reward value is +1; if the candidate answer is irrelevant to the corresponding cultural norm, the reward value is 0. Finally, strategy optimization feedback is performed, specifically, the reward values ​​are fed back to the model to achieve iterative optimization of the group-based relative strategy.

[0068] An evaluation was conducted on the technical solution presented in this example. The pre-selected questionnaires used in the evaluation were primarily based on three assessment sets: the World Values ​​Survey (WVS), the ValuesSurvey Module 2013 (VSM13), and Content Moderation (Mod.). Compared to existing methods based on cue engineering or data augmentation fine-tuning, the advantages of the technical solution presented in this example are: (1) Data costs are significantly reduced, and availability is strong in low-resource scenarios. Due to the use of a reinforcement learning mechanism of "standard pool + external reward model" instead of manually labeled answer labels, the model training does not rely on large-scale, balanced cultural label data; threshold filtering reduces irrelevant sample noise and ensures effective updates, thus achieving stable alignment benefits even in low-resource cultures. In Table 1, the bolded data represents the best-performing data under the same evaluation set and expected resource conditions.

[0069] Table 1. Comparison of cultural value alignment between regions with low-resource corpora and regions with high-resource corpora.

[0070] In this example, single-national culture reinforcement learning is a one-to-one model, while multi-national culture reinforcement learning is a one-to-many model. Since the evaluation metrics for the three evaluation sets differ, 1-JSD is used in the WVS dataset, Euclidean distance in the VSM13 dataset, and F1 score in the Mod. dataset. The F1 score is a statistical metric used to evaluate the performance of binary classification models, ranging from 0 to 1. JSD is the Jensen-Shannon divergence; a higher 1-JSD value indicates better cultural value alignment; a smaller Euclidean distance indicates better cultural value alignment; and a higher F1 score indicates better cultural value alignment. As shown in Table 1, the technical solution in this example maintains good cultural value alignment in both low-resource and high-resource corpus regions, demonstrating excellent cultural alignment stability.

[0071] (2) Generative tasks perform better, and cultural norms are more fully internalized. By using "normative cluster reward + relative advantage within the group (GRPO)," the judgment of "whether it conforms to the target cultural values" is directly written as a reward signal into the policy update. Compared to SFT (Supervised Fine-Tun-ing) / Prompt (cue words), which only imitates the sample distribution, it can maintain stronger consistency and decisiveness in open-ended generation. For example... Figure 3 As shown, this diagram illustrates the performance of the example technique (cultural reinforcement learning method) and the traditional supervised fine-tuning method on open-ended problems. The test results compare the performance of the example technique and the supervised fine-tuning method in open-ended generation across nine cultural contexts.

[0072] (3) Improved cross-cultural generalization and training stability. Relying on the "normative pool" of abstract concept-level clustering as the alignment anchor, the technical solution in this example achieves stable advantages in both "one-to-one" (single culture) and "one-to-many" (multicultural unification) training paradigms. Compared with other traditional methods, it performs better. Specific indicators are shown in Table 2. In Table 2, the bolded data represents the best results in the same cultural context. The meaning of the data is the same as in Table 1, and will not be repeated here.

[0073] Table 2. Comparison of the degree of alignment of cultural values ​​in different cultural contexts.

[0074] The cultural value question-answering model training method provided in this application determines the question cluster distance for each target cluster based on the question vector corresponding to the training question and the cluster center of the target cluster. The minimum value among the question cluster distances is taken as the minimum distance. If the minimum distance is less than or equal to a preset distance, the target cluster corresponding to the minimum distance is taken as the training cluster corresponding to the training question. If the minimum distance is greater than the preset distance, the training question is determined to be an irrelevant question to filter the training questions and avoid polluting the target question-answering model. Furthermore, for each candidate output, the consistency result is determined based on the candidate output and the training statement set, and the reward value corresponding to the candidate output is determined based on the consistency result. This facilitates the subsequent use of the group relative optimization strategy model to update the target question-answering model. This achieves thresholded sample filtering and the construction of a suitable reward-driven strategy, and performs reasonable intra-group relative advantage strategy optimization in cultural alignment, thereby improving the training and updating effect of the target question-answering model.

[0075] Figure 4 This is a schematic diagram of a cultural value question-and-answer model training device provided in an embodiment of this application. The device includes: an information extraction module 410, a reward value calculation module 420, and a model training and update module 430.

[0076] The information extraction module 410 is used to acquire training questions, determine training clusters corresponding to the training questions based on the question vectors corresponding to the training questions and the cluster centers of each target cluster, and determine a set of training statements from a hierarchical statement list based on the training clusters and the training countries corresponding to the training questions; wherein the target clusters and cluster centers are the clustering results of the question vectors corresponding to each survey question; the hierarchical statement list is a list of cultural value statements constructed using each country and each target cluster as indexes; the reward value calculation module 420 is used to determine each candidate output and its corresponding reward value based on the set of training statements, the target question-answering model, the preset number of question-answering attempts, the training questions, and the corresponding training countries; the model training and update module 430 is used to determine the relative reward advantage based on each reward value, and update the target question-answering model based on the training questions, each candidate output, and the relative reward advantage through a group relative optimization strategy model.

[0077] Optionally, it also includes: a basic information analysis module, used to determine each target cluster, the cluster center of each target cluster, and the hierarchical statement list based on the following methods: Based on the question vector corresponding to each survey question in the cultural value question and answer set, a clustering operation is performed to obtain each target cluster and its corresponding cluster center. Using each target cluster and each country in the cultural value question and answer set as an index, the cultural value question and answer set is processed to construct a hierarchical statement list. The cultural value question and answer set includes basic information about each survey question, the countries corresponding to each survey question, the frequency of responses, the mainstream answers, and the mainstream intensity for each country.

[0078] Optionally, before performing clustering operations based on the question vectors corresponding to each survey question in the cultural value question-and-answer set to obtain each target cluster and its corresponding cluster center, the basic information analysis module is further used to, for each survey question in the pre-selected questionnaire, determine the frequency of responses and the proportion of answer selections for each country on the survey question based on the survey answers of each country to the survey question; for each pair of countries to be calculated, determine the similarity of answers between the two countries on the survey question based on the proportion of answer selections for the survey question between the two countries; and determine the similarity of answers between the two countries on the survey question based on the proportion of answer selections for the survey question between the two countries. Based on the corresponding response frequencies, a similarity threshold is determined for the two countries to be calculated on the survey question. If the answer similarity exceeds the similarity threshold, the survey answers for the two countries to be calculated on the survey question are filtered. The response frequencies of each filtered country on the survey question are determined. Based on the basic information of the survey question, the filtered countries, the response frequencies of each filtered country, the mainstream answer corresponding to the maximum value in the answer selection ratio, and the mainstream strength, standardized question-and-answer data for the survey question is constructed. Based on the standardized question-and-answer data for each survey question in the pre-selected questionnaire, a cultural value question-and-answer set is constructed.

[0079] Optionally, the basic information analysis module is also used to determine the cultural value theme and theme confidence level corresponding to each survey question in the pre-selected questionnaire based on a preset large language model; in response to the theme confidence level being less than the preset confidence level, the survey question is deleted to update the pre-selected questionnaire.

[0080] Optionally, the reward value calculation module 420 is further configured to input the training question and the corresponding training country into the target question-answering model according to a preset number of question-answering attempts to obtain each candidate output; for each candidate output, determine each consistency result based on the candidate output and the training statement set, and determine the reward value corresponding to the candidate output based on each consistency result.

[0081] Optionally, the reward value calculation module 420 is further configured to, for each target question in the training statement set, perform consistency determination on the target value statement and the candidate output according to a preset large language model, and determine the consistency result corresponding to the target question; wherein, the target value statement corresponding to the target question includes the frequency of answers, mainstream answers, and mainstream strength of the training countries on the target question; determine whether there is at least one contradiction among the consistency results; if there is at least one contradiction, determine the reward value corresponding to the candidate output as a first value; if none of the consistency results are contradictory, determine whether there is at least one support among the consistency results; if there is at least one support, determine the reward value corresponding to the candidate output as a second value; if none of the consistency results are support, determine the reward value corresponding to the candidate output as a third value; wherein, the first value is less than the third value, and the third value is less than the second value.

[0082] Optionally, the information extraction module 410 is further configured to, for each target cluster, determine the problem cluster distance based on the problem vector corresponding to the training problem and the cluster center of the target cluster, and take the minimum value among the problem cluster distances as the minimum distance; in response to the minimum distance being less than or equal to a preset distance, take the target cluster corresponding to the minimum distance as the training cluster corresponding to the training problem; in response to the minimum distance being greater than the preset distance, determine the training problem as an irrelevant problem.

[0083] The cultural value question-and-answer model training device provided in this application can execute the cultural value question-and-answer model training method of any embodiment of this application and can have the same technical effect as the cultural value question-and-answer model training method in the above embodiments.

[0084] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 5 As shown, the electronic device 500 includes one or more processors 501 and memory 502.

[0085] The processor 501 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 500 to perform desired functions.

[0086] The memory 502 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 501 may execute the program instructions to implement the cultural value question-answering model training method of any embodiment of this application described above, and / or other desired functions. Various contents such as initial extrinsic parameters and thresholds may also be stored in the computer-readable storage medium.

[0087] In one example, the electronic device 500 may further include an input device 503 and an output device 504, these components being interconnected via a bus system and / or other forms of connection mechanisms (not shown). The input device 503 may include, for example, a keyboard, a mouse, etc. The output device 504 may output various information to the outside, including warning messages, braking force, etc. The output device 504 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0088] Of course, for the sake of simplicity, Figure 5 Only some of the components of the electronic device 500 relevant to this application are shown in this illustration; components such as buses, input / output interfaces, etc., are omitted. In addition, the electronic device 500 may include any other suitable components depending on the specific application.

[0089] In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps of the cultural value question-and-answer model training method provided in any embodiment of this application.

[0090] The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this application. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0091] Furthermore, embodiments of this application may also be computer-readable storage media storing computer program instructions, which, when executed by a processor, cause the processor to perform the steps of the cultural value question-and-answer model training method provided in any embodiment of this application.

[0092] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0093] It should be noted that the terminology used in this application is for the purpose of describing specific embodiments only and is not intended to limit the scope of this application. As shown in the specification and claims of this application, unless the context clearly indicates otherwise, words such as "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, or apparatus. Without further limitations, an element defined by the phrase "comprising an..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element.

[0094] It should also be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on this application. Unless otherwise expressly specified and limited, the terms "installed," "connected," "linked," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication between two elements. For those skilled in the art, the specific meaning of the above terms in this application can be understood according to the specific circumstances.

[0095] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. The above descriptions are only preferred embodiments of this application. It should be noted that due to the limitations of written expression, while there are objectively infinite specific structures, those skilled in the art can make several improvements, modifications, or changes without departing from the principles of this invention, and can also combine the above technical features in an appropriate manner. These improvements, modifications, changes, or combinations, or the direct application of the inventive concept and technical solution to other situations without modification, should all be considered within the scope of protection of this application.

Claims

1. A method for training a cultural value question-and-answer model, characterized in that, include: The training question is obtained, and the training cluster corresponding to the training question is determined based on the question vector corresponding to the training question and the cluster center of each target cluster. The training statement set is determined from the hierarchical statement list based on the training cluster and the training country corresponding to the training question. The target cluster and the cluster center are the clustering results of the question vectors corresponding to each survey question. The hierarchical statement list is a list of cultural value statements constructed with each country and each target cluster as indexes. Based on the training statement set, the target question-answering model, the preset number of question-answering attempts, the training question, and the corresponding training country, determine each candidate output and its corresponding reward value. Based on each reward value, a relative reward advantage is determined, and the target question-answering model is updated using a group relative optimization strategy model, based on the training question, each candidate output, and the relative reward advantage.

2. The method according to claim 1, characterized in that, Each target cluster, the cluster center of each target cluster, and the hierarchical statement list are determined based on the following method: Based on the question vector corresponding to each survey question in the cultural value question and answer set, a clustering operation is performed to obtain each target cluster and its corresponding cluster center; Using each target cluster and each country in the cultural value question and answer set as an index, the cultural value question and answer set is processed to construct a hierarchical statement list; The cultural value question and answer set includes basic information about each survey question, as well as the countries corresponding to each survey question, the frequency of responses, the mainstream answers, and the mainstream intensity for each country.

3. The method according to claim 2, characterized in that, Before performing clustering operations based on the question vector corresponding to each survey question in the cultural value question-and-answer set to obtain each target cluster and its corresponding cluster center, the following steps are also included: For each survey question in the pre-selected questionnaire, based on the survey answers from each country to the survey question, determine the corresponding response frequency and the answer selection ratio for each country on the survey question; For each pair of countries to be calculated, the similarity of their answers to the survey question is determined based on the proportion of their responses to the survey question. Based on the response frequencies of the two countries to be calculated on the survey question, a similarity threshold for the two countries to be calculated on the survey question is determined. If the similarity of the answers is greater than the similarity threshold, then the survey answers corresponding to the two countries to be calculated for the survey question are filtered. Determine the response frequency of each filtered country on the survey question. Based on the basic information of the survey question, the filtered countries, the response frequency of each filtered country, the mainstream answer corresponding to the maximum value in the answer selection ratio, and the mainstream strength, construct standardized question-and-answer data for the survey question. Based on the standardized data of the questions and answers for each survey question in the pre-selected questionnaire, a set of questions and answers on cultural values ​​is constructed.

4. The method according to claim 3, characterized in that, Also includes: For each survey question in the pre-selected questionnaire, the cultural value theme and theme confidence level corresponding to the survey question are determined based on a preset large language model. If the confidence level of the topic is lower than the preset confidence level, the survey question is deleted to update the pre-selected survey questionnaire.

5. The method according to claim 1, characterized in that, The step of determining each candidate output and its corresponding reward value based on the training statement set, the target question-answering model, the preset number of question-answering attempts, the training question, and the corresponding training country includes: According to the preset number of questions and answers, the training questions and the corresponding training countries are input into the target question-answering model to obtain each candidate output; For each candidate output, a consensus result is determined based on the candidate output and the training statement set, and the reward value corresponding to the candidate output is determined based on the consensus result.

6. The method according to claim 5, characterized in that, The step of determining each consistency result based on the candidate output and the training statement set, and determining the reward value corresponding to the candidate output based on each consistency result, includes: For each target question in the training statement set, the target value statement and the candidate output are evaluated for consistency based on a pre-defined large language model to determine the consistency result corresponding to the target question; wherein, the target value statement corresponding to the target question includes the frequency of responses, mainstream answers, and mainstream intensity of responses from the training countries to the target question; Determine whether at least one of the consistent results is contradictory; If at least one consistent result is contradictory, then the reward value corresponding to the candidate output is determined to be the first value; If all consistent results are not contradictory, then determine whether there is at least one supporting result among the consistent results; If at least one consistent result is supported, then the reward value corresponding to the candidate output is determined to be the second value; If none of the consistency results support the candidate output, then the reward value corresponding to the candidate output is determined to be the third value. Wherein, the first value is less than the third value, and the third value is less than the second value.

7. The method according to claim 1, characterized in that, The step of determining the training cluster corresponding to the training problem based on the problem vector corresponding to the training problem and the cluster center of each target cluster includes: For each target cluster, the cluster distance is determined based on the question vector corresponding to the training question and the cluster center of the target cluster, and the minimum value among the cluster distances is taken as the minimum distance. If the minimum distance is less than or equal to a preset distance, then the target cluster corresponding to the minimum distance is taken as the training cluster corresponding to the training problem; If the minimum distance is greater than the preset distance, then the training problem is determined to be an irrelevant problem.

8. A training device for a cultural value question-and-answer model, characterized in that, include: The information extraction module is used to obtain training questions, determine the training clusters corresponding to the training questions based on the question vectors corresponding to the training questions and the cluster centers of each target cluster, and determine the training statement set from the hierarchical statement list based on the training clusters and the training countries corresponding to the training questions; wherein, the target clusters and cluster centers are the clustering results of the question vectors corresponding to each survey question; the hierarchical statement list is a list of cultural value statements constructed with each country and each target cluster as indexes; The reward value calculation module is used to determine each candidate output and its corresponding reward value based on the training statement set, the target question answering model, the preset number of question answering times, the training question, and the corresponding training country. The model training and update module is used to determine the relative reward advantage based on each reward value, and update the target question answering model based on the training question, each candidate output, and the relative reward advantage through a group relative optimization strategy model.

9. An electronic device, characterized in that, The electronic device includes: Processor and memory; The processor executes the steps of the cultural value question-answering model training method as described in any one of claims 1 to 8 by calling the program or instructions stored in the memory.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program or instructions that cause a computer to perform the steps of the cultural value question-and-answer model training method as described in any one of claims 1 to 8.