Mahjong decision method based on unified action vector space and legal constraints

CN122702153APending Publication Date: 2026-09-08CHINA UNIV OF GEOSCIENCES (WUHAN) +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610835705.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-10
Publication Date
2026-09-08

AI Technical Summary

Technical Problem

[0007]本发明的目的在于:提出一种基于统一动作向量空间与合法约束的麻将决策方法,解决现有技术中麻将多类型动作表达割裂、非法动作干扰模型输出、不同地方麻将玩法下动作空间适配能力不足、合法动作约束层次单一以及规则切换与决策流程协同性较差的问题

Benefits of technology

1.本发明通过建立统一动作向量空间,将多类型麻将动作统一映射至同一输出空间,使得策略模型能够在同一向量结构中对不同类型动作进行评分与比较,避免了传统分离建模方式导致的模型结构碎片化问题,有利于提高训练流程和推理接口的一致性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122702153A_ABST
    Figure CN122702153A_ABST
Patent Text Reader

Abstract

The application relates to the mahjong artificial intelligence and intelligent decision-making technical field and discloses a mahjong decision-making method based on a unified action vector space and legal constraints, which comprises the following steps: obtaining a game state and constructing a feature representation; analyzing a rule configuration file, dynamically establishing a unified action vector space, and mapping various mahjong actions to unique indexes; generating a legal action constraint mask consistent with the action space dimension, and distinguishing forced, preferred and prohibited actions; a strategy model outputs an action score vector, determines a target action index after mask constraint processing, and performs reverse decoding and execution; when a rule condition is met or a configuration file is switched, the action space and constraint generation logic are updated, and a rule branch decision is executed. The application avoids model structure fragmentation through a unified action space, improves the rationality of legal action selection by using a layered mask, realizes multi-game dynamic adaptation with the aid of rule configuration driving, and enhances the generalization ability and expandability of the mahjong decision-making system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and intelligent decision-making technology in mahjong, and in particular to a mahjong decision-making method based on a unified action vector space and legal constraints. Background Technology

[0002] Mahjong is a typical discrete game task involving incomplete information, multiple stages, and multiple participants. Unlike games with complete information such as Go and chess, mahjong games contain a large amount of hidden information, and the set of actionable actions for each player at any given moment changes in real time with the state of the game. For example, during the round of drawing tiles, players typically engage in actions such as discarding, concealed kong, or supplementing a kong; when responding to other players' discards, they may engage in actions such as eating, ponging, exposed kong, winning, or passing. Furthermore, different regional variations of mahjong differ significantly in terms of the conditions for eating tiles, whether honor tiles can form sequences, the handling of bonus tiles or special tiles, and the rules governing winning hands and scoring patterns. This makes mahjong decision-making tasks complex in terms of rules, heterogeneous in actions, and diverse in scenarios.

[0003] Existing intelligent decision-making methods for mahjong games mainly suffer from the following problems: First, the expression of actions is often quite fragmented. Some methods model actions such as playing, eating, ponging, konging, and winning as different classification problems, or use multiple output heads and multiple sub-models for separate predictions. While such methods can accomplish decision-making tasks to some extent, the different output structures for different action types make it difficult for the model to directly compare multiple types of actions in a unified space. This easily leads to problems such as fragmented model structure, complicated training processes, and inconsistent inference interfaces.

[0004] Secondly, the constraints on legal actions are often not fully integrated into the model's reasoning process. The executable actions in a Mahjong game are not a fixed set, but rather determined by the current game state, rule conditions, and candidate hand combinations. If the model directly scores all actions without transforming the set of legal actions in the current situation into a constraint mechanism that matches the model's output structure, the model may assign higher scores to illegal actions, thus affecting the legality of the final decision and the system's stability. Especially in engineering deployment scenarios, if the agent outputs actions that do not satisfy the current situation's rule constraints, it can directly lead to system failure or the server refusing to execute.

[0005] Third, existing methods typically lack a unified rule management mechanism for adapting to different regional Mahjong variations. For different regional Mahjong variations, existing systems often require modifications to program logic, switching hard-coded rule branches, or even retraining the model for adaptation. This approach is not only costly to maintain but also hinders the unified deployment of multiple variations, making it difficult to meet the demands of real-world business scenarios for rapid switching between variations and flexible rule expansion.

[0006] Therefore, it is necessary to provide a new intelligent decision-making method and system for Mahjong. By constructing a unified action vector space, multiple types of Mahjong actions are mapped to the same output space; by constructing a legal action constraint mask with the same dimension as the unified action vector space, rule constraints are transformed into vector-level constraint calculations; and by using a rule engine driven by rule configuration files, dynamic adaptation to different local Mahjong gameplay is achieved, thereby improving the consistency, legality, real-time performance, and scalability of the decision-making process. Summary of the Invention

[0007] The purpose of this invention is to propose a mahjong decision-making method based on a unified action vector space and legal constraints, which solves the problems in the existing technology such as fragmented expression of various types of mahjong actions, interference of illegal actions with model output, insufficient adaptability of action space under different local mahjong gameplay, single level of legal action constraints, and poor coordination between rule switching and decision-making process.

[0008] Specifically, this invention provides a mahjong decision-making method based on a unified action vector space and legal constraints, the method comprising the following steps: S1. Obtain the current mahjong game status data, and perform feature encoding on the status data to construct a status feature representation; S2. Parse the currently loaded rule configuration file, extract the action type set and action parameters under the current mahjong gameplay, dynamically establish a unified action vector space based on the action type set, and map the various mahjong actions of the preset and rule definition to the unique index position in the unified action vector space. S3. Based on the current game state, rule conditions, and set of executable actions, generate a legal action constraint mask that is consistent with the dimension of the unified action vector space. The legal action constraint mask is used to distinguish between rule-mandated executable actions, strategy-prioritized actions, and rule-prohibited actions. S4. Input the state feature representation into a preset strategy model, and have the strategy model output the original action score vector in the unified action vector space; S5. Based on the legal action constraint mask, the original action score vector is constrained to suppress the score of the position corresponding to the rule-prohibited action, and the score of the position corresponding to the rule-mandated executable action and the policy-priority action is retained or weighted, and finally the target action index is determined in the legal action set. S6. Based on the target action index and its index range, perform action de-decoding to recover the corresponding actual mahjong operation and output the execution; S7. When the current game meets the preset rule conditions or the rule configuration file is switched, the dimension of the unified action vector space, the mapping relationship of the action index interval, and the generation logic of the legal action constraint mask are updated by reloading the corresponding rule configuration file, and the corresponding rule branch decision is executed.

[0009] A storage medium storing instructions and data for implementing a mahjong decision-making method based on a unified action vector space and legal constraints.

[0010] A mahjong decision-making device based on a unified action vector space and legal constraints includes: a processor and a storage medium; the processor loads and executes instructions and data in the storage medium to implement a mahjong decision-making method based on a unified action vector space and legal constraints.

[0011] The beneficial effects provided by this invention are: 1. This invention establishes a unified action vector space, mapping multiple types of mahjong actions to the same output space. This enables the strategy model to score and compare different types of actions within the same vector structure, avoiding the fragmentation problem of the model structure caused by traditional separate modeling methods. This is beneficial to improving the consistency of the training process and inference interface.

[0012] 2. This invention extracts the action type set and action parameters based on the rule configuration file, and dynamically adjusts the dimension and index interval mapping relationship of the unified action vector space, so that the action space can adapt to the changes in action types and the number of candidate actions under different local mahjong gameplay, thereby improving the system's scalability and adaptation efficiency in multi-game scenarios.

[0013] 3. This invention generates a legal action constraint mask with the same dimension as the unified action vector space, and distinguishes between rule-mandated executable actions, policy-prioritized actions, and rule-prohibited actions in the legal action constraint mask. This unifies the rule legality constraint and action priority constraint into a vector-level constraint calculation process, thereby effectively preventing illegal actions from entering the final candidate results and improving the rationality and stability of legal action selection.

[0014] 4. This invention organically combines action space mapping, legal action constraints, policy network scoring, and action anti-decoding process to form a complete decision-making chain of "state encoding - unified mapping - constraint scoring - index decoding", which has strong systematicity and engineering feasibility.

[0015] 5. By introducing rule configuration files and a rule engine mechanism, this invention can achieve the linkage update of action space mapping relationships, legal action constraint generation logic, and rule branch processing flow under different local Mahjong gameplays without modifying the main structure of the strategy model. This improves the maintainability, cross-game generalization ability, and engineering application value of the system. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 This is a schematic diagram of the unified action vector space and the legal action constraint process; Figure 3 This is a schematic diagram of the rule branch switching process driven by rule configuration; Figure 4 This is a schematic diagram of the hardware device operation according to an embodiment of the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0018] Before formally describing the present invention, a general description of the solution of the present invention will be given first to facilitate understanding.

[0019] Example 1 Please refer to Figures 1-3 The present invention provides a mahjong decision-making method based on a unified action vector space and legal constraints, comprising the following steps: S1. Obtain the current mahjong game status data, and perform feature encoding on the status data to construct a status feature representation; It should be noted that the state data includes at least the current decision-making player's hand information, their own revealed cards, their own discarded cards, other players' revealed cards, other players' discarded cards, the currently active card information, the current seat information, the information of the player initiating the current action, and rule parameter information. The state data is then structured and feature-encoded to form a state feature representation suitable for input to the strategy model.

[0020] S2. Parse the currently loaded rule configuration file, extract the action type set and action parameters under the current mahjong gameplay, dynamically establish a unified action vector space based on the action type set, and map the various mahjong actions of the preset and rule definition to the unique index position in the unified action vector space. It should be noted that establishing the unified action vector space in step S2 specifically includes: S21. Assign non-overlapping index ranges to pass, win, play, eat, pong, open kong, concealed kong, supplementary kong, and special actions defined by the rules. Among them, the play action is continuously indexed based on the card value, so that each playable card corresponds to a unique index position. S22. For the action of taking a card, based on the combination of the current card and the candidate straights, the candidate straights are sorted and combined with the relative position of the card in the straights to calculate a unique action index. S23. For actions such as pong, open kong, concealed kong, and supplementary kong, mapping is performed based on the index range corresponding to the tile value and the action type, thereby realizing the distinguishable representation of different kong types in a unified action vector space.

[0021] Specifically, the system parses the currently loaded rule configuration file, extracts the action type set and action parameters for the corresponding Mahjong gameplay, and establishes a unified action vector space based on the action type set. It maps actions such as passing, winning, discarding, eating, ponging, open kong, concealed kong, supplementary kong, and special actions defined in the rules to unique index positions in the unified action vector space, enabling the strategy model to output an action score vector with dimensions consistent with the unified action vector space.

[0022] In this invention, the unified action vector space is preferably constructed using a segmented index mapping method, with each action type occupying a different index interval, and different candidate operations under the same type of action corresponding to different index positions within the index interval; when the rule configuration file is switched, causing changes in the action type set, action parameters, or special action types, the dimension and index interval mapping relationship of the unified action vector space can be adjusted accordingly.

[0023] Furthermore, assuming the currently loaded rule configuration file is R, the total dimension of the unified action vector space can be expressed as:

[0024] in, These represent the dimensions corresponding to the actions of passing, winning, and playing cards, respectively. These represent the action dimensions corresponding to drawing, ponging, open kong, concealed kong, and supplementary kong, respectively. For use with rule configuration files The corresponding action enable indicator value is 1 when the corresponding action type is allowed in the current gameplay, and 0 otherwise. This represents a special action range dimension dynamically defined by the rules configuration file, used to accommodate special card exchanges, wild card replenishment, special card swaps, or other custom actions for gameplay.

[0025] For any action type k, the system assigns a corresponding index range. The index range satisfies:

[0026]

[0027]

[0028] in, This indicates the starting index of the index range corresponding to action type k. Indicates the termination index. This indicates the interval dimension corresponding to the action type under the current rule configuration file. Through the above recursive method, action intervals can be reallocated when switching rule configuration files, thereby achieving hot updates and dynamic expansion of the unified action vector space.

[0029] For any candidate action 'a', the system executes the action encoding function according to its action type and action parameters. This yields the unique index of the candidate action in the current unified action vector space:

[0030] in, Indicates the action type. This indicates the card value, sequence combination, kong type, or special action parameter corresponding to the action. Therefore, variations in action types, the number of candidate actions, and the expansion of special actions across different regional Mahjong variations can all be uniformly mapped to the same index system.

[0031] As one embodiment, in this embodiment, the system establishes a unified action vector space to uniformly represent different types of actions in Mahjong. The unified action vector space is preferably a one-dimensional high-dimensional sparse vector space, whose dimension is determined by the set of action types defined in the current rule configuration file, the number of candidate actions corresponding to each action type, and special action parameters. Unlike a fixed-length action vector space, the unified action vector space in this embodiment is not pre-fixed for a single gameplay mode, but is constructed based on the currently loaded rule configuration file to adapt to the differences in action types and the number of candidate actions under different regional Mahjong gameplay modes.

[0032] Specifically, the system first reads the current rule configuration file, extracting whether eating cards is allowed under the current gameplay, whether open kongs are allowed, whether concealed kongs or supplementary kongs exist, whether special action types are defined, and rule parameters related to action coding. Then, based on the action type set, it assigns corresponding index intervals to different action types and establishes a mapping relationship between action types and index intervals. In this way, when gameplay rules change causing changes in the action type set, the system can reconstruct a unified action vector space and synchronously update the index intervals corresponding to each action type, thereby ensuring that the action coding structure remains consistent with the current gameplay rules.

[0033] To ensure the distinguishability of different types of actions within a unified space, this embodiment preferably employs a segmented index mapping method to construct a unified action vector space. Specifically, the action of passing a card is mapped to a fixed index position, and the action of winning a card is mapped to another fixed index position; the action of discarding a card is mapped to a continuous index interval, ensuring that each card value corresponds to a unique index; the action of taking a card is mapped to a dedicated interval, used to express multiple candidate actions for taking cards corresponding to different suits, different sequence structures, and different operation card positions; the actions of ponging, open kong, concealed kong, and supplementary kong each occupy their respective index intervals and are mapped to a unique position within the corresponding interval according to the card value; when special action types are defined in the rule configuration file, corresponding extended index intervals are further allocated for the special action types.

[0034] For card-playing actions, this embodiment maps each playable card value to a consecutive index position. For actions like "pong" and "kong," a unique action index is obtained based on the combination relationship between card value and action type, using the method of "card value index + action interval offset." This method has the advantages of stable indexing, simple implementation, and easy alignment with the output structure of neural networks.

[0035] For the "eat" action, since it depends not only on the target card value but also on the straight combination structure and the relative position of the card in the straight, a finer-grained unique index mapping method is used. Specifically, the system generates candidate straight combinations based on the current card, sorts the candidate straights, and calculates its unique index position in the "eat" interval by combining the position and suit of the card in the straight. In this way, different "eat" combinations can correspond to different indices in a unified action vector space, thus achieving unique encoding and unique inverse encoding of the "eat" action.

[0036] In a preferred implementation, after completing the action range allocation, the system further generates a mapping table of "action type - index range - encoding rule". This mapping table is used not only to encode candidate actions in the current situation into a unified action vector space, but also to recover the actual mahjong operation based on the target action index in the subsequent action de-decoding stage. When the rule configuration file is switched, the system regenerates the mapping table and uses the updated mapping table as the basis for subsequent legal action constraint mask generation and action de-decoding. Therefore, rule switching no longer only affects rule branch judgment, but also synchronously affects the construction process of the unified action vector space.

[0037] Through the above design, the system can uniformly map all executable actions under the current gameplay to a single action scoring vector. This eliminates the need for the neural network model to construct different output heads for different action types. Instead, it only needs to output a scoring vector with the same dimension as the current unified action vector space to complete unified inference for all action types. This design not only improves the consistency of the model's output structure but also avoids model structure fragmentation caused by differences in action types and gameplay switching, thereby enhancing the system's scalability and engineering adaptability in various mahjong scenarios.

[0038] S3. Based on the current game state, rule conditions, and set of executable actions, generate a legal action constraint mask that is consistent with the dimension of the unified action vector space. The legal action constraint mask is used to distinguish between rule-mandated executable actions, strategy-prioritized actions, and rule-prohibited actions. It should be noted that generating the legal action constraint mask in step S3 specifically includes: S31. For any action index in the unified action vector space Define the mandatory legal mask respectively. Policy priority mask and forced illegal masking The value of each mask is either 0 or 1; S32. Configure the forced valid mask, policy priority mask, and forced invalid mask to satisfy a mutual exclusion or weak mutual exclusion relationship:

[0039] S33. Based on the current game state and rule conditions, assign corresponding mask values ​​to each action index position to form a hierarchical legal action constraint mask that is consistent with the dimension of the unified action vector space.

[0040] Specifically, based on the current game state, rule conditions, and set of executable actions, a legal action constraint mask with the same dimension as the unified action vector space is generated. This legal action constraint mask is used to represent the constraint levels of different actions in the current situation, including at least rule-mandated executable actions, strategy-prioritized actions, and rule-prohibited actions, thus providing a basis for action selection and scoring constraints in the subsequent reasoning stage.

[0041] Furthermore, the legal action constraint mask preferably adopts a hierarchical mask structure. For any action index 'a' in the unified action vector space, the mandatory legal mask, policy priority mask, and mandatory illegal mask are defined as follows:

[0042]

[0043]

[0044] in, Used to mark actions in the current situation that require the rules to retain decision-making eligibility. Used to mark actions that have high strategic value in the current hand state. Used to mark actions that are prohibited under the current rule conditions.

[0045] The three types of masks mentioned above satisfy a mutual exclusion or weak mutual exclusion relationship, that is:

[0046] For other action positions that are not encoded as mandatory legal actions or policy-priority actions, a default invalidation flag or a prohibition flag can be assigned according to implementation requirements. In this way, the system can not only express whether an action is legal, but also further express the differences in rule enforceability and policy priority between different legal actions.

[0047] As one embodiment, in this embodiment, to ensure that the reasoning results on the unified action vector space meet the constraints of the current game rules and to further reflect the priority differences between different legal actions, the system constructs a hierarchical legal action constraint mask based on the unified action vector space. The hierarchical legal action constraint mask is no longer limited to a binary judgment of "legal / illegal" action execution, but is at least divided into three constraint levels: rule-mandated executable actions, strategy-prioritized actions, and rule-prohibited actions.

[0048] Specifically, for any action index 'a' in the unified action vector space, the system defines a mandatory legality mask. Policy priority mask and forced illegal masking And satisfy: , ,

[0049] The above three types of masks satisfy a mutual exclusion or weak mutual exclusion relationship:

[0050] Among them, the mandatory legal mask is used to mark actions that must retain decision-making eligibility under the current rules, such as actions that must be played after drawing a tile, actions that must enter the winning candidate when the winning conditions are met, and actions that can only be selected from the declared actions after entering a specific declaration stage; the strategy priority mask is used to mark actions that are not the only mandatory actions but have high strategic value in the current game, such as actions that lead to a winning hand after a kong, actions that shorten the number of points to a winning hand, actions that increase the number of effective tiles drawn, and actions that balance the benefits of the winning hand with risk control; the mandatory illegal mask is used to mark actions that are explicitly prohibited by the current rules, such as eating tiles that are not allowed by the rules, special kong types that are not supported by the rules, actions that cannot be established under the current candidate hand set, and other actions that do not appear in the current legal candidate set.

[0051] During the generation process, the system first reads the current game state, rule configuration file, and current set of executable actions; then, based on the constructed "action type - index range - encoding rule" mapping table, it encodes each candidate action into a corresponding index position in a unified action vector space; finally, it combines rule conditions, hand type evaluation results, and strategy evaluation results to assign corresponding hierarchical mask values ​​to each index position.

[0052] Specifically, for strategy-priority actions, their priority scores can be pre-calculated using a card efficiency algorithm. Preferably, let the strategy priority of candidate action a be... Therefore, this priority can be calculated by comprehensively considering factors such as changes in the number of tiles to be played, changes in the number of effective tiles drawn, changes in the winning hand pattern, changes in risk costs, and changes in the value of special tiles. Thus, the layered mask not only expresses whether an action is allowed to be executed, but also expresses the priority relationship between allowed actions.

[0053] S4. Input the state feature representation into a preset strategy model, and have the strategy model output the original action score vector in the unified action vector space; It should be noted that the state feature representation obtained in step S1 is input into the policy model, and the output is an action score vector in a unified action vector space. The policy model is preferably a deep neural network model, and more preferably a convolutional neural network model or a policy network model containing a convolutional feature extraction structure.

[0054] S5. Based on the legal action constraint mask, the original action score vector is constrained to suppress the score of the position corresponding to the rule-prohibited action, and the score of the position corresponding to the rule-mandated executable action and the policy-priority action is retained or weighted, and finally the target action index is determined in the legal action set. It should be noted that determining the target action index in step S5 specifically includes: S51. Obtain the original action score vector output by the strategy model. And the strategy priority score pre-calculated from card efficiency analysis. ; S52. The original scores are fused based on the layered mask and policy priority to calculate the final action score:

[0055] in , , For the preset weighting coefficients, and This is a penalty constant used to suppress actions prohibited by the rules; S53, From the current set of candidate actions Select the action index with the highest final score as the target action index: .

[0056] As one embodiment, the action score vector is constrained with the legal action constraint mask, the score of the position corresponding to the rule-prohibited action is suppressed, the score of the position corresponding to the rule-mandated executable action and the policy-priority action is retained or weighted, and the target action index with the best score is selected from the legal action set as the current decision output.

[0057] In this way, the legal action constraint mask is not only used to exclude illegal actions, but also to reflect the constraint hierarchy and priority relationship of different actions within the set of legal actions.

[0058] In a preferred implementation, let the original action score vector output by the strategy model be... Then, the fused action score based on the layered mask can be expressed as:

[0059] in, For the preset weighting coefficients, and Ideally, the penalty constant should be large enough to suppress actions prohibited by the rules; This represents the strategy priority score for candidate action a, which can be calculated from card efficiency analysis, changes in the number of tiles to be ready, effective number of tiles drawn, hand pattern benefits, risk assessment, or a combination thereof.

[0060] Based on the fused action score, the system determines the target action index as follows:

[0061] Here, A represents the set of candidate actions in the current unified action vector space. Subsequently, the system determines the target action index... The mapping relationship between the index range and the current rule configuration file is used to perform anti-decoding of the action to obtain the corresponding actual mahjong action.

[0062] In this way, rule constraints and strategy preferences are considered together in the unified scoring process, so that the final output action not only meets the legality requirements of the rules, but also has higher decision rationality and stability.

[0063] S6. Based on the target action index and its index range, perform action de-decoding to recover the corresponding actual mahjong operation and output the execution; Specifically, based on the action range to which the target action index belongs and the corresponding index mapping relationship, the target action index is decoded into actual mahjong operations, including passing, winning, playing specific card values, eating specific sequence combinations, ponging, open kong, concealed kong, supplementary kong, or special actions, and output to the upper-level business system or game environment for execution.

[0064] It should be noted that the action de-decoding in step S6 specifically includes: S61. Based on the target action index The index range to which it belongs Determine the corresponding action type; S62. Call the anti-decoding function corresponding to the action type. By combining the encoding parameters in the current rule configuration file with the current set of candidate actions, a class-by-class matching is performed to recover the actual mahjong operation:

[0065] S63. Perform consistency verification on the recovery result according to the current rule configuration file. If the verification passes, output the execution. If the verification fails, reselect the suboptimal action according to the preset rollback strategy or output the default legal action.

[0066] As one embodiment, in this embodiment, after completing the unified action scoring and hierarchical legal action constraint processing, the system selects the final action index from the current unified action vector space and restores the actual mahjong action based on the index interval mapping relationship. Unlike the method of only performing maximum value selection on the original model output, the action selection process in this embodiment is based on "dynamic expansion of unified action space + hierarchical mask constraint fusion", thus it can simultaneously adapt to the changes in action sets and differences in action priorities in different local mahjong gameplay.

[0067] Specifically, the system first obtains the unified action score vector output by the strategy model, and then generates the fused action score according to the hierarchical legal action constraint mask described in Example 3. Then, the optimal action is selected from the current candidate action set A:

[0068] in, This represents the target action index. Since the candidate action set has already been encoded based on the current rule configuration file, the current game state, and the set of executable actions, the target action index... It is consistent with both the current gameplay rules and the feasibility of actions in the current situation. After obtaining the target action index... Then, the system performs action de-decoding. Preferably, the de-decoding process includes the following steps: First, based on the range to which the target action index belongs. Determine its corresponding action type; Secondly, call the de-decoding function corresponding to this action type. ( The system combines the encoding parameters in the current rule configuration file with the current set of candidate actions to perform a class-by-class matching. Finally, the corresponding actual mahjong game operation is restored. This process can be formally represented as:

[0069] in, This is the action de-decoding function corresponding to the rule configuration file R. The actual actions of the mahjong tiles that are finally restored.

[0070] For playing a card, the system restores the specific card value based on the index offset within the playing range; for actions like ponging, open kong, concealed kong, and supplementary kong, the system restores the corresponding card set based on the action type and card value parameters; for taking a card, the system restores the specific taking combination based on the currently played card, candidate sequence combinations, and the relative position of the played card within the sequence; for special actions, the system restores the corresponding action content based on the special action parameters, trigger conditions, and output format defined in the rule configuration file. Therefore, even if the total dimension of the unified action vector space, the position of the action range, and the structure of special actions change when the rule configuration file is switched, the system can still complete the action restoration based on the updated encoding and decoding mappings.

[0071] In a preferred implementation, after performing action de-decoding, the system also performs a consistency check on the recovery result according to the current rule configuration file to confirm that the recovered action meets the action restrictions, special card handling conditions, and candidate action range restrictions in the current gameplay. If the check passes, the corresponding mahjong operation is output to the upper-level business system or the game environment for execution; if the check fails, a suboptimal action is reselected according to a preset rollback strategy, or a default legal action is output. This approach further improves the stability and engineering availability of action output in cross-gameplay deployment scenarios.

[0072] S7. When the current game meets the preset rule conditions or the rule configuration file is switched, the dimension of the unified action vector space, the mapping relationship of the action index interval, and the generation logic of the legal action constraint mask are updated by reloading the corresponding rule configuration file, and the corresponding rule branch decision is executed.

[0073] It should be noted that the update decision logic in step S7 specifically includes: S711. When a rule configuration file switch is detected, the action type set and action parameters under the new gameplay are re-parsed, and the total dimension of the updated unified action vector space is dynamically calculated:

[0074] in This is the enable indicator for the rule corresponding to the action type; S712. Based on the updated total dimension, reallocate index ranges for each action type and generate an updated action type-index range-encoding rule mapping table; S713. Based on the updated mapping table, synchronously update the generation logic of the legal action constraint mask, so that the action space mapping, constraint generation and rule branch processing can be coordinated and adapted.

[0075] It should be noted that the execution of the rule branch decision in step S7 specifically includes: S721. When the currently operated card is a special card or a treasure card, or when the current rule configuration file defines a special action condition that should be processed by the rule engine first, trigger the rule branch switch. In a preferred implementation of Leping Mahjong, the triggering conditions for rule branches include at least one of the following situations: the currently operated tile is empty; there is a special tile or a bonus tile and the currently operated tile is the same as the special tile or a bonus tile; the currently operated tile is a character tile and the current rule allows character tiles to form a sequence; or the current rule configuration file defines special action conditions that need to be processed by the rule engine first.

[0076] S722. Load the corresponding rule configuration file to build the rule engine. The rule engine generates auxiliary or alternative decision results based on card type splitting, head-to-head number analysis, hand type evaluation and special card value determination. In one implementation, the rule engine performs card type decomposition, count-to-read calculation, special card determination, executable action filtering, and action priority judgment on the current situation based on the card type parameters, action parameters, and special rule parameters defined in the rule configuration file.

[0077] S723. Use the results output by the rule engine as the final decision output, or use them as an auxiliary basis for the construction of the unified action vector space and the generation of legal action constraint masks.

[0078] Specifically, the results obtained from the above process can be used directly as the final decision output, and also as an auxiliary basis for constructing the unified action vector space and generating legal action constraint masks. For example, when there are special action types in the current gameplay, the rule engine can first determine the triggering conditions and candidate range of the special action according to the rule configuration file, and add the special action to the corresponding index interval of the unified action vector space; when a mandatory priority or prohibition condition is set for a certain type of action in the current gameplay, the rule engine can simultaneously adjust the constraint level of the corresponding action position in the legal action constraint mask. In this way, a linkage relationship is formed between the rule engine, the unified action encoding mechanism, and the legal action constraint mechanism.

[0079] As one embodiment, in order to support dynamic adaptation to Leping Mahjong and other local Mahjong gameplay, the system adopts a rule branch switching mechanism based on rule configuration files.

[0080] Unlike implementations that rely solely on hard-coded branch switching of local rules, the rule configuration file in this embodiment is used not only to build the rule engine but also to determine the foundation for constructing the unified action vector space under the current gameplay, the mapping relationship of action index intervals, and the generation logic of legal action constraint masks, thereby achieving coordinated updates of rule switching and decision-making processes.

[0081] It should be noted that step S0 is included before step S1, and step S0 includes: S01. Multiple sets of rule configuration files are pre-set using structured data format. Each rule configuration file is used to define key rule parameters in a mahjong game. The rule parameters include at least: action type set, executable action set constraints, whether eating is allowed, whether specific kong types are allowed, whether honor tiles can form a sequence, whether there are bonus tiles or special tiles, winning hand rules, and special action types. As one embodiment, specifically, the system pre-configures multiple sets of rule configuration files, which are preferably stored in JSON, XML, or other structured data formats. Each rule configuration file defines key rule parameters for a specific Mahjong gameplay mode. These rule parameters include at least: a set of action types, constraints on the set of executable actions, whether eating (chow) is allowed, whether specific kong (kong) types are allowed, whether honor tiles can form sequences, whether there are bonus tiles or special tiles, winning hand rules, special tile handling rules, special action types, and other parameters that affect the legality of actions and the action encoding method.

[0082] S02. When the system starts, it reads the specified rule configuration file, dynamically builds the rule engine based on the rule configuration file, and simultaneously determines the dimension of the unified action vector space under the current gameplay and the index interval mapping relationship corresponding to each action type.

[0083] When the system starts, it first reads the specified rule configuration file, dynamically builds the rule engine based on the rule configuration file, and simultaneously determines the dimension of the unified action vector space under the current gameplay and the index interval mapping relationship corresponding to each action type.

[0084] In a preferred implementation, after loading the rule configuration file, the system first parses the allowed action types and the number of candidate actions for each action type under the current gameplay. Then, based on the parsing results, it updates the structure of the unified action vector space, assigning corresponding index intervals to actions such as passing, winning, discarding, eating, ponging, open kong, concealed kong, supplementary kong, and special actions defined by the rules. After updating the action interval mapping relationship, the system further updates the generation logic of the legal action constraint mask based on the current rule conditions and the current game state, ensuring that subsequent action constraint masks maintain dimensional and semantic consistency with the updated unified action vector space. Therefore, the switching of the rule configuration file no longer only affects the rule engine itself, but simultaneously affects the action space construction, action constraint generation, and subsequent action selection processes.

[0085] When the system detects that the conditions are met, it can switch to the rule engine branch first. The rule engine will make supplementary or alternative decisions based on card type splitting, forward hand analysis, hand type evaluation, special card value determination and remaining card information. If the rule branch triggering conditions are not met, the system will continue to execute the strategy model decision process based on the unified action vector space and legal action constraint mask.

[0086] Through the above mechanism, when the game platform switches to another regional Mahjong gameplay, only the corresponding rule configuration file needs to be switched to complete the synchronous update of the rule engine, unified action vector space, action index interval mapping relationship, and legal action constraint mask generation logic, without modifying the core strategy model's main structure, rewriting the action coding logic, or retraining the model. This design expands the rule configuration-driven adaptation mechanism beyond the rule branches themselves to cover the complete collaborative process of "rule parsing—action space update—constraint generation—rule branch processing," thereby significantly enhancing the system's scalability, maintainability, consistency, and cross-gameplay adaptability.

[0087] Example 2: Please see Figure 4 , Figure 4 This is a schematic diagram of the hardware device in operation according to an embodiment of the present invention. The hardware device specifically includes: a mahjong decision-making device 401 based on a unified action vector space and legal constraints, a processor 402, and a storage medium 403.

[0088] A mahjong decision-making device 401 based on a unified action vector space and legal constraints: The mahjong decision-making device 401 based on a unified action vector space and legal constraints implements the mahjong decision-making method based on a unified action vector space and legal constraints.

[0089] Processor 402: The processor 402 loads and executes the instructions and data in the storage medium 403 to implement the Mahjong decision-making method based on a unified action vector space and legal constraints.

[0090] Storage medium 403: The storage medium 403 stores instructions and data; the storage medium 403 is used to implement the Mahjong decision-making method based on a unified action vector space and legal constraints.

[0091] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A Mahjong decision-making method based on a unified action vector space and legal constraints, characterized in that: Includes the following steps: S1. Obtain the current mahjong game status data, and perform feature encoding on the status data to construct a status feature representation; S2. Parse the currently loaded rule configuration file, extract the action type set and action parameters under the current mahjong gameplay, dynamically establish a unified action vector space based on the action type set, and map the various mahjong actions of the preset and rule definition to the unique index position in the unified action vector space. S3. Based on the current game state, rule conditions, and set of executable actions, generate a legal action constraint mask that is consistent with the dimension of the unified action vector space. The legal action constraint mask is used to distinguish between rule-mandated executable actions, strategy-prioritized actions, and rule-prohibited actions. S4. Input the state feature representation into a preset strategy model, and have the strategy model output the original action score vector in the unified action vector space; S5. Based on the legal action constraint mask, the original action score vector is constrained to suppress the score of the position corresponding to the rule-prohibited action, and the score of the position corresponding to the rule-mandated executable action and the policy-priority action is retained or weighted, and finally the target action index is determined in the legal action set. S6. Based on the target action index and its index range, perform action de-decoding to recover the corresponding actual mahjong operation and output the execution; S7. When the current game meets the preset rule conditions or the rule configuration file is switched, the dimension of the unified action vector space, the mapping relationship of the action index interval, and the generation logic of the legal action constraint mask are updated by reloading the corresponding rule configuration file, and the corresponding rule branch decision is executed.

2. The Mahjong decision-making method based on a unified action vector space and legal constraints as described in claim 1, characterized in that: The establishment of the unified action vector space in step S2 specifically includes: S21. Assign non-overlapping index ranges to pass, win, play, eat, pong, open kong, concealed kong, supplementary kong, and special actions defined by the rules. Among them, the play action is continuously indexed based on the card value, so that each playable card corresponds to a unique index position. S22. For the action of taking a card, based on the combination of the current card and the candidate straights, the candidate straights are sorted and combined with the relative position of the card in the straights to calculate a unique action index. S23. For actions such as pong, open kong, concealed kong, and supplementary kong, mapping is performed based on the index range corresponding to the tile value and the action type, thereby realizing the distinguishable representation of different kong types in a unified action vector space.

3. The Mahjong decision-making method based on a unified action vector space and legal constraints as described in claim 1, characterized in that, Step S3, generating the legal action constraint mask, specifically includes: S31. For any action index in the unified action vector space Define the mandatory legal mask respectively. Policy priority mask and forced illegal masking The values ​​of each mask are either 0 or 1; S32. Configure the forced valid mask, policy priority mask, and forced invalid mask to satisfy a mutual exclusion or weak mutual exclusion relationship: S33. Based on the current game state and rule conditions, assign corresponding mask values ​​to each action index position to form a hierarchical legal action constraint mask that is consistent with the dimension of the unified action vector space.

4. The Mahjong decision-making method based on a unified action vector space and legal constraints as described in claim 1, characterized in that, Step S5, determining the target action index, specifically includes: S51. Obtain the original action score vector output by the strategy model. And the strategy priority score pre-calculated from card efficiency analysis. ; S52. The original scores are fused based on the layered mask and policy priority to calculate the final action score: in , , For the preset weighting coefficients, and This is a penalty constant used to suppress actions prohibited by the rules; S53, From the current set of candidate actions Select the action index with the highest final score as the target action index: .

5. The Mahjong decision-making method based on a unified action vector space and legal constraints as described in claim 1, characterized in that: Step S6, specifically the action de-decoding, includes: S61. Based on the target action index The index range to which it belongs Determine the corresponding action type; S62. Call the anti-decoding function corresponding to the action type. By combining the encoding parameters in the current rule configuration file with the current set of candidate actions, a class-by-class matching is performed to recover the actual mahjong operation: S63. Perform consistency verification on the recovery result according to the current rule configuration file. If the verification passes, output the execution. If the verification fails, reselect the suboptimal action according to the preset rollback strategy or output the default legal action.

6. The Mahjong decision-making method based on a unified action vector space and legal constraints as described in claim 1, characterized in that: The specific steps in step S7 to update the decision logic include: S711. When a rule configuration file switch is detected, the action type set and action parameters under the new gameplay are re-parsed, and the total dimension of the updated unified action vector space is dynamically calculated: in This is the enable indicator for the rule corresponding to the action type; S712. Based on the updated total dimension, reallocate index ranges for each action type and generate an updated action type-index range-encoding rule mapping table; S713. Based on the updated mapping table, synchronously update the generation logic of the legal action constraint mask, so that the action space mapping, constraint generation and rule branch processing can be coordinated and adapted.

7. The Mahjong decision-making method based on a unified action vector space and legal constraints as described in claim 1, characterized in that: Step S7, which involves executing rule-based branch decisions, specifically includes: S721. When the currently operated card is a special card or a treasure card, or when the current rule configuration file defines a special action condition that should be processed by the rule engine first, trigger the rule branch switch. S722. Load the corresponding rule configuration file to build the rule engine. The rule engine generates auxiliary or alternative decision results based on card type splitting, head-to-head number analysis, hand type evaluation and special card value determination. S723. Use the results output by the rule engine as the final decision output, or use them as an auxiliary basis for the construction of the unified action vector space and the generation of legal action constraint masks.

8. The Mahjong decision-making method based on a unified action vector space and legal constraints as described in claim 1, characterized in that: Step S0 precedes step S1, and step S0 includes: S01. Multiple sets of rule configuration files are pre-set using structured data format. Each rule configuration file is used to define key rule parameters in a mahjong game. The rule parameters include at least: action type set, executable action set constraints, whether eating is allowed, whether specific kong types are allowed, whether honor tiles can form a sequence, whether there are bonus tiles or special tiles, winning hand rules, and special action types. S02. When the system starts, it reads the specified rule configuration file, dynamically builds the rule engine based on the rule configuration file, and simultaneously determines the dimension of the unified action vector space under the current gameplay and the index interval mapping relationship corresponding to each action type.

9. A storage medium, characterized in that: The storage medium stores instructions and data to implement the mahjong decision-making method based on a unified action vector space and legal constraints as described in any one of claims 1 to 8.

10. A mahjong decision-making device based on a unified action vector space and legal constraints, characterized in that: include: A processor and a storage medium; the processor loads and executes instructions and data in the storage medium to implement the mahjong decision-making method based on a unified action vector space and legal constraints as described in any one of claims 1 to 8.