Method and device for AI-based user matching

KR103022070B1Active Publication Date: 2026-09-21전준영
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
KR1020250173618
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2026-09-21
Estimated Expiration
2045-11-17

Smart Images

  • Figure 112025128321086-PAT00001_ABST
    Figure 112025128321086-PAT00001_ABST
Patent Text Reader

Abstract

The present disclosure discloses a method and a device comprising: a step of acquiring multiple modal data including profile text data, image data, and behavioral pattern data for multiple users; a step of generating multiple modal feature vectors including text feature vectors, image feature vectors, and behavioral feature vectors by inputting the text data, image data, and behavioral pattern data into corresponding artificial intelligence models based on the multiple modal data; a step of mapping the generated multiple modal feature vectors into a single integrated vector space through a multimodal fusion network; a step of calculating similarity between users within the single integrated vector space to calculate a matching score; and a step of collecting matching results and user feedback data in real time based on the matching score, and using the user feedback data for reinforcement learning to update the weights of the artificial intelligence model or the matching model in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The technical field of the present disclosure relates to a user matching method and device for maximizing user matching accuracy and satisfaction using artificial intelligence technology, and more specifically, to a method for providing suitable user matching information by integrating and analyzing multiple data types. Background Technology

[0002] With the recent popularization of smartphones, online and mobile-based matching services are being actively utilized. Existing matching services have primarily adopted heuristic or rule-based methods that calculate similarity based on static profile information (age, education, occupation, etc.) directly entered by users and simple survey responses. However, these existing technologies have the following limitations. First, they fail to fully utilize unstructured data containing potential preferences, such as user profile text (introduction) or images (photos), resulting in only superficial matching. Second, they lack the ability to improve models by reflecting actual user behaviors (profile browsing, continuing chats, accepting or rejecting matches, etc.) in real time during the matching process; consequently, matching accuracy remains fixed, or the potential for performance improvement based on accumulated data is limited. In other words, there was a problem in that they could not keep up with users' complex and changing preferences over time. Prior art literature

[0003] Korean Registered Patent No. 10-2159952 (September 25, 2020) Device and method for providing blind date arrangement service The problem to be solved

[0004] The problem to be solved by the present disclosure is to provide a more suitable user matching method by capturing in-depth and potential preferences through the analysis of user profile text, images, and behavioral pattern data within the service. means of solving the problem

[0005] As a technical means for achieving the technical problem described above, the device according to the first aspect of the present disclosure and the artificial intelligence-based user matching method may include: a step of acquiring multiple modal data including profile text data, image data, and behavioral pattern data for a plurality of users; a step of generating multiple modal feature vectors including text feature vectors, image feature vectors, and behavioral feature vectors by inputting the text data, image data, and behavioral pattern data into corresponding artificial intelligence models based on the multiple modal data; a step of mapping the generated multiple modal feature vectors into a single integrated vector space through a multimodal fusion network; a step of calculating a matching score by calculating similarity between users within the single integrated vector space; and a step of determining a matching result based on the matching score and providing it to a user terminal, and collecting user feedback data in real time and using the user feedback data for reinforcement learning, thereby updating the weights of the artificial intelligence model or the matching model in real time.

[0006] In addition, the above multiple modal data may further include chat data, and the behavioral pattern data may include user interaction indicators derived by analyzing the chat data.

[0007] In addition, the artificial intelligence model includes an artificial intelligence model for text processing, an artificial intelligence model for image processing, and an artificial intelligence model for behavioral pattern analysis, and the step of generating the multiple modal feature vectors may include a large-scale language embedding model specialized for Korean for the text processing artificial intelligence model to generate semantic embedding vectors from user profile sentences to generate the text feature vectors, and the artificial intelligence model for image processing may include a vision-language integration model to generate image embedding vectors representing visual tendencies from user profile images to generate the image feature vectors.

[0008] In addition, in the step of generating the multiple modal feature vectors, the artificial intelligence model for behavioral pattern analysis can generate the behavioral feature vectors by receiving the user's in-app activity logs, conversation frequency, response latency, and movement pattern data as input.

[0009] In addition, the multimodal fusion network includes an attention mechanism or a transformer-based fusion module, and the step of mapping to the single integration vector space can map to the single integration vector space by generating an optimal integration vector by assigning differential weights to feature vectors among the multiple modal feature vectors that have a high contribution to the matching accuracy.

[0010] Additionally, the step of calculating the similarity to produce a matching score may include a step of calculating the similarity using a hybrid calculation method combining cosine similarity and weighted Euclidean distance; and may further include a step of modeling user relationships within the single integrated vector space as a graph neural network and predicting potential favorability and relationship development possibilities between the users.

[0011] An artificial intelligence-based user matching device according to a second aspect of the present disclosure may include: a receiver that acquires multiple modal data including profile text data, image data, and behavioral pattern data for multiple users; and a processor that inputs the text data, the image data, and the behavioral pattern data into corresponding artificial intelligence models based on the multiple modal data to generate multiple modal feature vectors including text feature vectors, image feature vectors, and behavioral feature vectors, maps the generated multiple modal feature vectors into a single integrated vector space through a multimodal fusion network, calculates a matching score by calculating similarity between users within the single integrated vector space, determines a matching result based on the matching score and provides it to a user terminal, and collects user feedback data in real time and updates the weights of the artificial intelligence model or the matching model in real time by using the user feedback data for reinforcement learning. Effects of the invention

[0012] According to one embodiment of the present disclosure, by using multiple modal data such as text, images, and behavioral patterns to comprehensively identify user characteristics, it is possible to successfully perform matching between users with high potential intimacy rather than superficial matching.

[0013] In addition, by utilizing user feedback (actual matching success / failure, whether to continue chatting, etc.) in reinforcement learning, a reinforcement learning-based self-evolving system can be implemented in which matching results change and improve in real time according to user satisfaction.

[0014] Furthermore, improving matching accuracy can directly lead to user satisfaction and bring about key commercial effects, such as increasing service reuse rates and paid conversion rates.

[0015] The effects of the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by a person skilled in the art from the description below. Brief explanation of the drawing

[0016] FIG. 1 is a block diagram schematically illustrating the configuration of a device according to one embodiment. FIG. 2 is a flowchart illustrating each step of operation of a device according to one embodiment. FIG. 3 is a block diagram illustrating the overall processing flow of an artificial intelligence-based user matching method using multimodal feature fusion and reinforcement learning feedback according to one embodiment. FIG. 4 is a conceptual diagram schematically illustrating the process in which multiple modal feature vectors are mapped into a single integrated vector space and similarity between users is calculated according to one embodiment. FIG. 5 is a diagram schematically illustrating an example of a service home screen and a matching result screen provided to a user by a device according to one embodiment. Specific details for implementing the invention

[0017] The advantages and features of the present disclosure and the methods for achieving them will become clear by referring to the embodiments described below in detail together with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below but can be implemented in various different forms, and the embodiments provided are merely to make the disclosure complete and to fully inform those skilled in the art of the scope of the present disclosure.

[0018] The terms used in this specification are for describing embodiments and are not intended to limit the disclosure. In this specification, the singular form includes the plural form unless specifically stated otherwise in the text. The terms “comprises” and / or “comprising” as used in this specification do not exclude the presence or addition of one or more other components in addition to the components mentioned. Throughout the specification, the same reference numerals refer to the same components, and “and / or” includes each of the mentioned components and all combinations of one or more. Although terms such as “first,” “second,” etc., are used to describe various components, these components are not limited by these terms. These terms are used merely to distinguish one component from another. Accordingly, the first component mentioned below may be the second component within the technical scope of this disclosure.

[0019] Unless otherwise defined, all terms used herein (including technical and scientific terms) may be used in a meaning commonly understood by a person skilled in the art. Additionally, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise.

[0020] Spatially relative terms such as "below," "beneath," "lower," "above," and "upper" may be used to facilitate the description of the relationship between one component and other components as illustrated in the drawings. Spatially relative terms should be understood as encompassing different orientations of components during use or operation, in addition to the orientations depicted in the drawings. For example, if a component depicted in a drawing is inverted, a component described as "below" or "beneath" of another component may be placed "above" of that component. Therefore, the exemplary term "below" may encompass both the lower and upper directions. Components may also be oriented in other directions, and accordingly, spatially relative terms may be interpreted according to the orientation.

[0021] Various embodiments are described in detail below with reference to the drawings.

[0023] FIG. 1 is a block diagram schematically illustrating the configuration of a device (100) according to one embodiment.

[0024] Referring to FIG. 1, the device (100) may include a receiver (110) and a processor (120).

[0025] A receiving unit (110) according to one embodiment can acquire multiple modal data including profile text data, image data, and behavior pattern data for multiple users.

[0026] A processor (120) according to one embodiment can generate a multiple modal feature vector including a text feature vector, an image feature vector, and a behavior feature vector by inputting text data, image data, and behavior pattern data into corresponding artificial intelligence models based on multiple modal data. Additionally, the processor (120) can map the generated multiple modal feature vectors into a single integrated vector space through a multimodal fusion network. Additionally, the processor (120) can calculate a matching score by calculating the similarity between users within the single integrated vector space. Furthermore, the processor (120) can determine a matching result based on the matching score and provide it to a user terminal, and can update the weights of the artificial intelligence model or the matching model in real time by collecting user feedback data in real time and using the user feedback data for reinforcement learning. In one embodiment, the processor (120) can provide the matching result to the user terminal by controlling a transmitter (not shown).

[0027] In addition, it will be understood by those skilled in the art that other general components may be included in the device (100) in addition to the components illustrated in FIG. 1. The device (100) may further include a transmitter (not shown) that provides matching results and user matching information, and a memory (not shown) that stores multiple modal characteristic vectors, matching scores, feedback data, user matching information, etc. Alternatively, it will be understood by those skilled in the art that, according to other embodiments, some components among those illustrated in FIG. 1 may be omitted. In one embodiment, “data” may represent individual values, and “information” may be a concept representing a set of data according to a certain purpose.

[0028] A device (100) according to one embodiment can be used by a user and can be connected to all types of handheld-based wireless communication devices equipped with a touch screen panel, such as mobile phones, smartphones, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), tablet PCs, etc., and can also be included in or connected to devices that have a foundation for installing and running applications, such as desktop PCs, tablet PCs, laptop PCs, and IPTVs including set-top boxes.

[0029] The device (100) may be implemented as a terminal such as a computer that operates through a computer program to realize the functions described in this specification.

[0030] A device (100) according to one embodiment may include, but is not limited to, an artificial intelligence-based user matching system (not shown) and a related server (not shown). A server according to one embodiment may support an application that provides a user matching system.

[0031] In the following description, the focus will be on an embodiment in which a device (100) according to one embodiment acquires user matching information, but as previously mentioned, this may also be performed through interaction with a server. That is, the device (100) according to one embodiment and the server may be implemented as an integrated unit in terms of their functions, the server may be omitted, and it can be seen that this is not limited to any one embodiment.

[0032] In one embodiment, the device (100) and the server may be interconnected, and as information is transmitted and received via mobile communication with a user terminal or an expert terminal, the configuration of the device (100) may be performed by the server or by the device (100). For example, the device (100) may operate as a device or a server, and below, it will be described uniformly as the device (100).

[0034] FIG. 2 is a flowchart illustrating each step of operation of a device (100) according to one embodiment.

[0035] Referring to step S210, a device (100) according to one embodiment can acquire multimodal data including profile text data, image data, and behavior pattern data for a plurality of users.

[0036] In one embodiment, multiple users may be subjects using an application (hereinafter referred to as the “matching application”) that provides an AI-based user matching service supported by a device (100) according to one embodiment. The device (100) may receive various forms of information from each user through the matching application or automatically collect data from the user’s usage behavior. More specifically, profile text data may include self-introduction sentences, tastes and interests, preferred types, and descriptive phrases related to lifestyle that the user inputs when creating their matching profile. For example, if the user inputs sentences such as “I like reading books at cafes,” “I enjoy outdoor activities,” or “I prefer calm people,” such text data may be collected as key information reflecting the user’s personality, tastes, and relationship preferences. Additionally, the system may collect additional linguistic expression data as a supplement through chat windows or survey responses within the service. Furthermore, image data may be photo data provided by the user through a profile settings screen or gallery upload function within the application, and may include information expressing the user’s visual characteristics, such as profile photos, photos of hobbies, and travel photos. For example, the background, location, attire, facial expression, and composition of photos uploaded by a user can be utilized as elements reflecting that user's lifestyle or tendencies. This image data can be stored and managed so that it can be used for visual feature analysis in a later stage. Additionally, behavioral pattern data refers to non-verbal data automatically accumulated during the user's use of the app, and may include, for instance, the user's access frequency, activity time zones, patterns of accepting and rejecting matching requests, conversation duration, message response speed, location-based movement trends, and click history for specific content.These behavioral pattern data can indirectly reflect unstructured characteristics such as the user's social tendencies, activeness, and ability to maintain likeability. In one embodiment, the profile text data may include self-introduction and detailed information directly entered by the user during the profile creation process, as well as additional text included in external content disclosed by the user only if the user consents. Examples of external content include text data posted on social media profiles, personal blogs, or other websites. By analyzing this expanded text data to understand the user's tendencies in a more three-dimensional way, the matching accuracy can be improved. Additionally, multiple modal data may further include chat data, and the behavioral pattern data may include user interaction indicators derived by analyzing the chat data. For example, the device (100) can analyze conversation logs between users to calculate the diversity of conversation topics, frequency of message exchange, emotional response patterns, duration of conversation, and mutual response rate, and these indicators can be used as variables to quantitatively represent the mutual affinity, conversation immersion, or tendency to form likeability between users. The interaction indicator derived in this way reflects behavioral characteristics beyond the content of simple text messages, so it can be usefully utilized for reinforcement learning of the matching model or adjustment of personalized weights in the future. Accordingly, the device (100) can collect multiple modal data including text, images, chat, and behavioral data in real time through a user terminal or server network, and step S210 can function as a step of configuring an input base for an artificial intelligence model to analyze data and produce a matching result. In one embodiment, the user may include both men and women, and in the matching relationship, when a match is performed between a man and a woman, it can be performed as an example of a blind date.

[0037] Referring to step S220, a device (100) according to one embodiment can generate multiple modal feature vectors including text feature vectors, image feature vectors, and behavior feature vectors by inputting text data, image data, and behavior pattern data into corresponding artificial intelligence models based on multiple modal data. Step S220 can function as a modality-specific representation learning step that precedes the subsequent feature fusion and similarity calculation steps.

[0038] In one embodiment, the artificial intelligence model may include an artificial intelligence model for text processing, an artificial intelligence model for image processing, and an artificial intelligence model for behavioral pattern analysis. The artificial intelligence model for text processing may include a large-scale language embedding model specialized for the Korean language to generate semantic embedding vectors from user profile sentences and generate text feature vectors. Specifically, the artificial intelligence model for text processing may include a large-scale language embedding model specialized for the Korean language and may generate text feature vectors by producing semantic embedding vectors after preprocessing (normalization, tokenization, etc.) language data such as user profile self-introductions, interest keywords, survey responses, and (optional) chat text. For example, by applying a pre-trained model suitable for Korean and multilingual embeddings, the meaning of descriptive sentences such as “prefers quiet cafes” or “enjoys hiking activities” can be represented as low-dimensional or high-dimensional (e.g., hundreds to thousands of dimensions) vectors. When chat data is associated with text modals, the model may generate embeddings that reflect linguistic features such as conversation topics, tone of expression, and vocabulary diversity. In one embodiment, an artificial intelligence model for text processing can calculate linguistic characteristic indicators for each sentence or utterance unit, such as emotional tone (e.g., positive, negative, neutral), empathy keywords (e.g., frequency and pattern of empathy expressions like “I relate,” “me too,” etc.), linguistic distance (e.g., overlap of interest keywords between two users, similarity of value-related expressions), and ratio of positive / negative expressions. At this time, the semantic vector of each sentence and the linguistic characteristic indicators can be integrated into a single multidimensional vector, and, for example, can be normalized into a fixed-length vector of 768 dimensions or more to form a text characteristic vector.In one embodiment, the device (100) may apply a weighted average or attention-based aggregation method based on chronological order, conversational context, and utterance importance to aggregate sentence-unit embeddings into user-unit profile embeddings, and in this process, each sentence embedding may be scaled through L2 normalization or softmax-based weights to stably reflect the overall linguistic tendency while preventing excessive bias toward specific expressions.

[0039] Additionally, an artificial intelligence model for image processing may include a vision-language integration model to generate image feature vectors by generating image embedding vectors representing visual tendencies from user profile images. For example, the device (100) may construct image feature vectors by generating image embedding vectors representing visual tendencies from user profile photos or lifestyle photos. Specifically, by inferring visual cues such as background (indoor / outdoor), type of activity (hiking, cafe, etc.), and atmosphere (dynamic / static) within the photo and expressing them as quantified features, a vector in a form that can be utilized for integrated representation and similarity calculation in subsequent steps may be provided.

[0040] Additionally, the artificial intelligence model for analyzing behavioral patterns can generate the behavioral characteristic vector by receiving input data such as the user's in-app activity log, conversation frequency, response latency, and movement pattern. For example, the device (100) can generate the behavioral characteristic vector by receiving input time-series / categorical signals such as the user's in-app activity log, conversation frequency, response latency, history of accepting and rejecting matching requests, time zones of access, and location-based movement trends. Interaction indicators between users derived from chat data (e.g., mutual response rate, conversation duration, emotional response patterns, etc.) can be included as features of the behavioral modal. The model can output a vector representation that internalizes the user's social tendencies and interaction tendencies through statistical aggregation, time-series encoding, or graph / attention-based representation learning.

[0041] Each modality characteristic vector can be managed in the storage of the device (100) along with the modal identifier and timestamp, and personal identification information (name, contact information, etc.) can be stored separately and anonymized from the analysis embeddings to satisfy personal information protection and security requirements.

[0042] In one embodiment, the device (100) may be configured to calculate interaction indicators from chat data, and to dynamically apply different levels of importance to each indicator depending on the user state and service context. For example, in the case of new users in the initial stage of registration, since preferences have not yet stabilized, the system may give higher weight to relatively observable short-term signals. Specifically, indicators reflecting immediacy and responsiveness, such as the frequency of message exchange within a certain period, the response rate to the other party's message, and whether the conversation ends after the first message, may be evaluated first. On the other hand, in the case of users who have used the service for a long time and have sufficiently accumulated text, images, and behavioral characteristics, the weighting may be redistributed to place greater emphasis on qualitative aspects of conversation, such as diversity of topics, consistency of emotional expression, and continuity of long conversations. Additionally, if the service is operated in a hyper-local environment based on actual meetings within a radius of 30 km, the device (100) may further increase the weight of conversation continuity and mutual responsiveness (behavioral pattern data) for pairs of users grouped in the same region or the same community (e.g., same university email verification). For example, in matching users within the same campus, the degree of interruption in the conversation flow and the stability of mutual responsiveness—factors that influence the success of an actual meeting—are given greater importance; thus, the rhythm, continuity, and mutual responsiveness of the conversation can possess greater explanatory power than profile text or image data. Conversely, in scenarios involving long-distance or different communities, higher weight can be assigned to text data because the likelihood of an immediate offline meeting is low. Such differential weighting based on situational context—such as new / long-term, proximity / distance, or same / different communities—is defined policy-wise, allowing the matching process to incorporate actual signals of relationship formation that are difficult to capture through simple averages or uniform scoring.

[0043] As another example, the device (100) can obtain information regarding exam periods, festival periods, vacation periods, etc. For example, during the exam period of a college student service, simultaneous access and quick responses become difficult, so a low favorability rating may be misjudged based solely on "response delay." During this period, the weight of 'response speed' included in the behavioral pattern data can be lowered, and instead, the weight of text empathy signals (similar interests, matching conversation topics) included in the text data can be increased. Conversely, during periods when the demand for meetings increases, such as festivals or the beginning of vacations, the behavioral weight (access pattern, conversation continuity) for the behavioral pattern data can be adjusted upward. During late-night hours, there are many non-simultaneous conversations, so the noise of behavioral signals increases; therefore, the text and image weights can be temporarily increased to stabilize the prediction.

[0044] Referring to step S230, the device (100) according to one embodiment can map the generated multiple modal feature vectors into a single integrated vector space through a multimodal fusion network.

[0045] More specifically, the multimodal fusion network may include a neural network structure for learning semantic relationships between different modalities (e.g., text, images, behavioral patterns, etc.) and deriving common expressions from them. In one embodiment, the multimodal fusion network may include an attention mechanism or a transformer-based fusion module. Additionally, the device (100) may generate an optimal integration vector by assigning differential weights to feature vectors among multiple modal feature vectors that contribute significantly to the matching accuracy, and map this to a single integration vector space. Specifically, the attention mechanism may be configured to dynamically calculate the importance of each of the multiple modal feature vectors contributing to the overall matching judgment, and to assign greater weights to feature vectors that have a high influence on the matching accuracy. In one embodiment, the device (100) can calculate a similarity score s_i between the feature vectors (text vector v_text, image vector v_img, behavior vector v_beh) for each modality and a matching-related query vector q, and calculate a modal-specific weight α_i by applying softmax normalization to s_i. For example, using the score s_i for modal i, α_i can be defined as α_i = exp(s_i) / Σexp(s_j), and the final integrated vector V_int can be calculated in the form V_int = Σα_i · v_i. By applying softmax-based weight distribution in this way, the sum of all modal weights is normalized to 1, thereby preventing the phenomenon where the importance of a specific modal becomes excessively large, and at the same time, the weights can be adjusted in a continuous and differentiable manner according to changes in data quality and feedback results.

[0046] In another embodiment, the multimodal fusion network may operate according to the following pseudocode. A quality metric q_i (e.g., missing data, resolution, text length) can be calculated for each modal using a text vector v_text, an image vector v_img, and an action vector v_beh as input data. The device (100) can calculate a modal score s_i = f(v_i, q_i). Additionally, the device (100) can calculate a weight α_i by applying softmax(s_i). The device (100) can calculate V_int = Σα_i · v_i and use it as a single fusion vector, where the function f can be implemented as a linear combination, a multilayer perceptron (MLP), or a transformer-based subnetwork, etc.

[0047] Additionally, in one embodiment, when a specific user's text expression reveals a distinct tendency while the image features contain neutral information, the attention module can generate an optimal integrated vector by assigning a relatively high weight to the text feature vector. Furthermore, the transformer-based fusion module can learn the interdependence between each modality using a self-attention structure. At this time, the correlation between modals is calculated through Query, Key, and Value vectors, and through this, multiple modal feature vectors can be reconstructed into mutually complementary expressions in a single semantic space. The device (100) can generate an optimal Integrated Feature Vector for each user through the fusion process, and this integrated vector can be mapped to a point within a Single Unified Embedding Space. The Single Unified Embedding Space is a high-dimensional semantic space that numerically represents the user's comprehensive tendency combining text, image, and behavioral features, and can serve as the basis for similarity calculation and matching score calculation performed in subsequent steps. In addition, in one embodiment, the multimodal fusion network may be configured to maintain or improve overall matching performance even if the importance of a specific modal changes over time by dynamically adjusting weights for each modality using user feedback or matching result data during the learning process.

[0048] In one embodiment, the device (100) can acquire multiple modal data based on weights that are gradually lowered in the order of behavioral pattern data, profile text data, and image data. For example, in blind dates / matching, since the "interaction actually continued" ultimately has the greatest predictive power, behavioral patterns (response interval, conversation maintenance, acceptance / rejection patterns, etc.) can be reflected first, followed by text (self-introduction, preferences, consistency / variety of conversational vocabulary), and finally images (profile, lifestyle context) can be reflected as auxiliary signals. This has the effect of avoiding image-centric temporary biases and leading to stable matching centered on behavioral signals that are more direct to actual likeability formation.

[0049] In one embodiment, the device (100) can collect user feedback (such as 'likes', conversation maintenance, reporting / blocking, etc.) in real time and gradually adjust the modal weights (text / image / behavior). Specifically, if it is observed that feedback from a recent period has consistently made a positive contribution to the prediction of a specific modal, the system may be designed to gradually increase the weight of that modal to quickly capture changes in the preferences of the recent user group. Conversely, if a certain modal is noisy or unreliable (e.g., low-resolution images, failure to detect faces, suspected auto-generated text, missing logs) and exceeds a certain threshold, a rule may be established to automatically and gradually reduce the contribution of that modal during that period and gradually return it to its original level when the data quality is restored. If necessary, the total sum of the modal weights may always be maintained within a constant range, and a stabilization device may be included to set minimum and maximum limits for individual modals to prevent sudden modal collapse or over-reliance on a single modal. This adjustment maintains a moderate rate of change so as not to blindly follow only ultra-short-term feedback, and can prevent overfitting through a buffering mechanism that reflects the flow of accumulated feedback.

[0050] In one embodiment, the multimodal fusion network may be configured to automatically calculate the importance of each modal using an attention mechanism, while considering input quality signals and situational context together in the calculation process. For example, in the case of an image, if reliability is evaluated as low due to resolution, face detection, suspicion of excessive correction or synthesis, or absence of metadata, the attention mechanism may visibly reduce the contribution of the image and relatively increase the weight of text or behavioral modals. In the case of text, if it is determined that the information density is low because the sentence length is extremely short—for example, shorter than a preset length—or the repetition of the same sentence is more frequent than a preset number of times, the system may slightly reduce the importance of the text while compensatorily increasing the weight of behavioral language signals, such as topic diversity and conversational continuity observed in the chat. Additionally, situational context may also be reflected. For example, during periods when the likelihood of offline meetings is high, such as daytime on weekends, behavioral modals (matching activity times, rhythm of responses, and actual movement patterns) are treated as relatively more important. Conversely, during periods with frequent non-simultaneous access, such as late-night hours or exam periods, static information from images and text demonstrates more stable predictive power; therefore, attention can increase the weight of these two modals for a certain period. Given the scarcity of chat and behavioral logs for new users, multiple modal data can be acquired based on weights assigned to image data, profile text data, and behavioral pattern data, which are gradually reduced in that order. In other words, a phased priority switching policy may be included, which prioritizes the use of information from image and profile text data and gradually increases behavioral and chat-based signals once conversations accumulate beyond a certain threshold.

[0051] Referring to step S240, a device (100) according to one embodiment can calculate a matching score by calculating the similarity between user vectors mapped to a single integrated vector space in step S230.

[0052] More specifically, the device (100) can quantitatively evaluate the propensity similarity between two users by measuring the distance or directional agreement between each user vector within a single integrated vector space. In one embodiment, the device (100) can calculate a matching score that simultaneously reflects semantic similarity and numerical proximity by applying a hybrid calculation method that combines Cosine Similarity and Weighted Euclidean Distance. For example, Cosine Similarity represents the directional agreement of two user propensities based on the angle between vectors, and Euclidean Distance can express the quantitative difference in propensity by reflecting the difference in distance between vectors. The device (100) can calculate a matching score that balances various user characteristics (e.g., propensity, image, behavioral pattern, etc.) by adjusting weights based on the relative importance of these two similarity measures. Additionally, in one embodiment, the device (100) may calculate a ranking of suitability among users based on a calculated matching score, and if it is above a certain threshold value, classify it as a matching candidate group or display it as a recommendation target. In one embodiment, the device (100) may first select the top N candidates (e.g., top 20, 50) for each user based on a matching score calculated in a single integrated vector space, and then construct a final recommendation list by reflecting exposure limit rules to prevent excessive exposure of the same candidate (e.g., maximum number of exposures within a certain period), reliability correction based on recent feedback (e.g., score attenuation of candidates that have been repeatedly reported and blocked), and user-specific search and utilization ratios (recommendation of new candidates for search vs. repeated recommendation of stable candidates).At this time, the device (100) may additionally apply diversity constraints to prevent excessive bias in gender, age group, region, interest group, etc. within the recommendation list, thereby enabling the user to explore a wide range of candidates. The matching score can be corrected by reinforcement learning through user feedback in a subsequent step (S250) and used to continuously improve the matching accuracy of the entire system.

[0053] In one embodiment, the device (100) may, instead of using the cosine similarity and distance-based indicators calculated in a single integrated vector space as is, preemptively correct for differences in distribution at the cohort level (region, university, grade, etc.) to ensure fairness and consistency in matching. Since there may be a tendency for vectors to cluster in a specific direction among user groups with the same characteristics, if normalization of the mean and variance of the corresponding cohort is performed first and then similarity is calculated to reduce recommendation bias, overestimation due to regional and group effects may be mitigated. Additionally, some users may have an abnormally large number of images or excessively long text, causing dimensional skewness; in this case, the system may gradually readjust the importance of each dimension of the fused vector according to the history of training data and feedback, thereby preventing a specific dimension from irrationally dominating the overall similarity. Rapid changes in similarity caused by temporary noise (e.g., event-driven profile changes, one-time conversation surges) can be reflected gradually through buffering rules, and safety can be ensured by reflecting them conservatively in cases accompanied by negative feedback, such as reports or blocking.

[0054] Referring to step S250, the device (100) according to one embodiment determines a matching result based on a matching score and provides it to a user terminal, and can update the weights of an artificial intelligence model or a matching model in real time by collecting user feedback data in real time and using the user feedback data for reinforcement learning.

[0055] More specifically, the matching model can improve the accuracy of future matching judgments by using a matching score calculated within a single integrated vector space and user responses (e.g., 'like', 'rejection', conversation duration, etc.) as reward signals, and dynamically adjusting weight parameters between each feature vector. In one embodiment, the reward function may be defined not as a simple binary reward for a single event, but as a weighted sum of multiple behavioral indicators. For example, the device (100) may calculate a primary reward (α) for whether a mutual 'like' is successful, an additional reward (β) when the conversation duration is maintained for longer than a preset threshold time, a reward (γ) when the mutual response rate exceeds a threshold ratio, and a penalty (δ) for negative feedback such as reporting, blocking, or immediate exit, and define the weighted average of these values ​​as a final reward value R = α·w₁ + β·w₂ + γ·w₃ - δ·w₄. Here, w₁, w₂, w₃, and w₄ are weighting coefficients that can be adjusted according to service policy or learning results, and the device (100) can strengthen or weaken the influence of indicators that are evaluated as more important in relationship formation through these coefficients.

[0056] In one embodiment, an artificial intelligence model (e.g., a model for text processing, image processing, or behavior analysis) basically performs representation learning to extract feature vectors from user data. Although it is not the primary target for updating in reinforcement learning, if the amount of accumulated feedback reaches a threshold or a certain period, some weights may be updated using a transfer learning or fine-tuning method. That is, the artificial intelligence model may be a learning target at the representation level to maintain or correct the quality of the input data representation, and the matching model may be a learning target at the judgment level that adapts in real time based on user feedback. The two models can be interconnected to achieve continuous improvement in matching accuracy. By configuring it in this way, the device (100) can gradually improve the reliability and personalization level of the matching results and implement a self-evolving matching system through repeated user interaction.

[0057] In one embodiment, the device (100) may operate the update cycles of the matching model and the artificial intelligence model (embedding generation model) separately. While the matching model immediately receives user feedback and adjusts weights in minutes to hours, the artificial intelligence model that generates text, images, and behavior embeddings may be configured to perform batch retraining on a daily or weekly basis for overall service stability, or to perform lightweight updates on only parts using parameter efficiency techniques (e.g., adapter / LoRa, etc.). Specifically, the matching model may perform parameter updates in mini-batch units by triggering when feedback on new matching results accumulates to a predetermined number (e.g., 10 cases, 50 cases, etc.) or when a predetermined time interval (e.g., 10 minutes, 1 hour, etc.) has elapsed. At this time, the device (100) may achieve a balance between real-time adaptability and learning stability by combining an online learning mode that uses only feedback from the previous period and a batch and semi-online mode that performs stable updates using feedback accumulated over a certain period (e.g., 1 day, 1 week). This separate operation has the advantage of minimizing unexpected quality fluctuations caused by frequent software updates by simultaneously securing the agility of real-time personalization and the stability of the presentation model. In addition, updates to the embedding model are performed only when explicit triggers occur, such as the detection of service metric degradation or data distribution shifts, thereby reducing unnecessary retraining costs.

[0058] Additionally, the device (100) may model user relationships within a single integrated vector space as a Graph Neural Network (GNN) and predict potential affinity and relationship development possibilities between users. The Graph Neural Network can learn structural patterns of the entire relationship network by representing each user as a node and similarity or interaction history between users as edges. Through this, the device (100) can predict potential affinity, relationship development possibilities, or long-term mutual preference possibilities between users, and can implement a relationship-centric prediction model that goes beyond simple vector similarity-based matching.

[0059] In one embodiment, the device (100) can go beyond comparing only points in a single integrated vector space and construct a relationship graph with users as nodes and interactions such as similarity, conversation history, and response rate as edges, and infer using a graph neural network (GNN). In this case, the system can consider time information together and reflect recency weighting so that interactions that occurred recently have a greater influence. Empirically considering that patterns with many common connections, such as patterns where three-way relationships (friends of friends) are frequently formed within the same community, the same club, or subjects, have a positive effect on actual matching success, the system can maintain the highest weight assigned to behavioral pattern data because behavior and relationship signals are more decisive than text / images in this case, while assigning an intermediate weight to text data to complement the qualitative interpretation of relationships (conversation tone, values), and assigning a relatively low weight to image data. In one embodiment, in the case of a cold start user, since sufficient chat data and behavioral pattern data do not yet exist, the device (100) may insert the user into a graph in an initial mixed state with representative vectors of the same community (e.g., same campus, same interest cluster, etc.). At this time, the system may apply a Relationship Prior to estimate user characteristics based on a weighting structure centered on image data initially, but configure the system so that the dominance of weighting gradually shifts to behavioral pattern data and text (relational conversation) data as time passes and actual conversation and behavioral data accumulates. That is, initially, visual and external signals account for a relatively high proportion, but as the user's actual interactions accumulate, the proportion of behavioral and relationship signals can be automatically strengthened.Since this procedure operates in a structure that mitigates the cold start problem while gradually reducing reliance on relationship priors as personal data accumulates, it can simultaneously improve recommendation accuracy and personalization levels in service environments with frequent influx of new users.

[0060] Additionally, the device (100) can collect matching results and user feedback data in real time and use them as reward signals for reinforcement learning. For example, if a user selects 'Like' or 'Reject' regarding a proposed matching result, or if the actual conversation duration lasts longer than a certain standard, these can be reflected as positive rewards. Based on such feedback data, the device (100) can continuously learn to reflect changes in the user's preference tendencies or social interaction patterns over time by updating the weights of the artificial intelligence model or the matching model in real time. Through this, the matching accuracy is gradually improved by repetitive user activity, and automatic enhancement of service quality can be achieved.

[0061] In one embodiment, the device (100) can design user responses to matching results as multi-level rewards. For example, a high reward can be given when a mutual 'like' is achieved, a medium reward if the conversation is maintained for a certain period of time or longer, and a small reward if the response rate exceeds a certain level, thereby capturing continuous positive events closely. Conversely, if the user leaves immediately after matching or reports and blocks, a strong penalty can be imposed to prevent the recommendation strategy from being repeated in similar situations later. In the reinforcement learning process, the search rate can always be maintained at a certain level to prevent the problem of learning stagnation caused by repeatedly using only already verified strategies. In addition, weighting adjustments based on personal information or sensitive attributes are prohibited by policy, and chat logs are used for analysis only with anonymized pattern indicators without storing the original text, thereby protecting user privacy. In the course of operation, fairness constraints can also be applied by monitoring cases where exposure or matching success rates for a specific cohort are excessively skewed, and automatically imposing a penalty if the deviation exceeds a set limit. This enables the simultaneous achievement of improved recommendation quality and responsible AI operations.

[0062] FIG. 3 is a block diagram illustrating the overall processing flow of an artificial intelligence-based user matching method using multimodal feature fusion and reinforcement learning feedback according to one embodiment.

[0063] Referring to FIG. 3, an AI-based user matching system according to one embodiment may include an input layer, a multi-modal processing layer, a feature fusion layer, and a real-time learning layer. The input layer may include user profiles, chat conversations, uploaded photos, and behavioral data obtained from multiple users. These data may be provided as multi-modal inputs required for subsequent AI analysis. In the multi-modal processing layer, an AI model corresponding to each data type may be applied. Specifically, a text processing unit may fine-tune a Korean-specific language model (e.g., Korean BERT) to generate a 768-dimensional text embedding vector that semantically vectorizes user self-introductions or conversation content. In addition, the Image Processing Unit can analyze user profile photos and lifestyle images using a vision-language integration model based on CLIP or Vision Transformer, and convert the results into 512-dimensional style vectors. For behavioral pattern data, vectors quantifying behavioral characteristics can be generated by analyzing log data such as user access frequency, response time, and matching acceptance / rejection history. The multiple modal feature vectors generated in this way can be combined in the Feature Fusion Layer. The Feature Fusion Layer may include Vector Concatenation and Weighted Fusion with Attention modules.In one embodiment, the attention mechanism dynamically calculates the relative importance of each modality (text, image, behavior) and assigns greater weight to modals that contribute more to matching accuracy. Through this, the device can generate an optimal Integrated Feature Vector for each user. This optimal Integrated Feature Vector serves as a point that numerically represents the user's overall propensity and can be mapped to a single integrated vector space. Next, a Hybrid Similarity Calculation Module can calculate the similarity between users within the single integrated vector space. In one embodiment, the similarity calculation is performed using a hybrid method combining Cosine Similarity and Weighted Euclidean Distance, which allows for the simultaneous consideration of directional match and distance-based proximity. The value calculated as a result of this calculation is defined as a Matching Score and can quantitatively represent the propensity-based fit between users. The calculated matching score is passed to the candidate sorting and recommendation stage by the matching algorithm, and a matching result can be generated. The matching result and user reactions (e.g., 'Like', conversation duration, response rate, etc.) can be fed back to the Real-Time Learning Layer. In the Real-Time Learning Layer, a Reinforcement Learning Feedback Loop operates to update the weights of the matching model using user feedback as a reward signal. Through this, the system evolves autonomously through repetitive user interactions, and matching accuracy can gradually improve over time.In addition, the system in one embodiment inputs the hybrid similarity calculation results into a deep learning-based matching algorithm (Deep-Tech AI Matching Platform) to finally output matching scores and recommendation results between users. Accordingly, as illustrated in FIG. 3, one embodiment can implement a continuously evolving AI-customized matching system that goes beyond simple profile-based matching by combining multimodal fusion technology that considers the complex nature of input data with a reinforcement learning loop based on user feedback. In one embodiment, Korean-BERT, CLIP, Vision Transformer, etc., may be used, but is not limited thereto. The embedding dimension can also be set in various ways depending on the implementation environment.

[0064] FIG. 4 is a conceptual diagram schematically illustrating the process in which multiple modal feature vectors are mapped into a single integrated vector space and similarity between users is calculated according to one embodiment.

[0065] Referring to FIG. 4, an AI-based user matching system according to one embodiment can integrate multiple modal feature vectors extracted from text, image, and behavioral data and map them to points within a single integrated vector space. As illustrated in the figure, each point represents an individual user, and different colors distinguish between user tendencies or interest groups (e.g., music enthusiasts, sports enthusiasts, readers, travelers, etc.). As illustrated in the lower left corner of FIG. 4, the evaluation of similarity between users can be performed by combining Cosine Similarity and a Distance Metric. Cosine Similarity is a method of measuring directional similarity based on the magnitude of the angle formed by two vectors; it has the characteristic that the value approaches 1 as the vectors are in the same direction and approaches 0 as they are in opposite directions. On the other hand, the Distance Metric may be a method of evaluating the propensity-based proximity of two users by calculating an actual numerical distance based on the positional difference between the two vectors. In other words, since the angle between the two vectors indicates the degree of agreement of the user's 'tendency direction' and the magnitude of the distance signifies the 'quantitative difference' of the tendency, a more precise matching score can be calculated by considering these two criteria together.

[0066] In the center of Fig. 4, an integrated vector of a new user is depicted, and this vector can receive a similarity score calculated based on its positional relationship with existing user groups. For example, a similarity value of 0.92 shown in the figure indicates that the new user has a high degree of propensity match with a specific user group (e.g., the reader group), while 0.78 may indicate a relatively low degree of match with another group (e.g., the traveler group). Such similarity scores can be displayed as matching lines, and the length or angle of each line can intuitively represent the distance and directional relationship between users. Additionally, a multi-modal vector configuration is depicted in the bottom right of the figure, where text, image, and behavioral vectors are represented in an overlapping form. This implies that characteristics extracted from different modalities are combined within a single space to represent user propensity in a multidimensional manner. For example, linguistic features (text) derived from a user's self-introduction sentence, visual features (images) extracted from a profile image, and usage behavior patterns (behaviors) are integrated so that each user can be mapped to unique location coordinates. Accordingly, FIG. 4 can be described as a conceptual diagram visually explaining the principle by which a system according to the present invention quantifies similarity between users based on a multimodal integrated vector space and calculates the pair of users with the highest similarity as matching candidates. With such a configuration, the matching system in one embodiment can reflect the user's internal tendencies and preferences more precisely than a simple keyword-based match determination.

[0067] FIG. 5 is a diagram schematically showing an example of a service home screen and a matching result screen provided to a user by a device (100) according to one embodiment.

[0068] Referring to FIG. 5, an AI-based user matching system according to one embodiment can visualize the matching results and display them on the screen of a user terminal. The screen on the left may show an example where a user is matched with a counterpart user determined to be an ideal type through the matching algorithm of the present invention. The screen may display the profile picture, age, and basic information of the matched counterpart user, and a unique matching number (e.g., 26,585) may be displayed in the center or at the top of the screen. The user may access the matching results through guidance messages such as “Check the profile of your ideal type and try chatting,” and may move to an actual conversation screen through a selection button such as “Would you like to connect?”. This configuration allows the matching results to be intuitively recognized and can induce the user to quickly start a conversation with the matched counterpart. The screen on the right of FIG. 5 illustrates an example of a relationship prediction and personality analysis screen provided in the post-matching stage. In one embodiment, the device (100) may quantitatively display the potential for relationship development and the degree of personality match between users, respectively, based on a matching score calculated through multimodal fusion and reinforcement learning. For example, quantified prediction results such as “Relationship Prediction 78%” or “Personality Analysis 92%” may be displayed on the screen; these values ​​may be calculated by synthesizing data analysis results, such as the similarity between the user’s propensity vectors, the consistency of behavioral patterns, and emotional response patterns. Additionally, the graph area illustrated in the diagram can visually represent how relationship indicators between the user and the match change over time or with feedback, allowing the user to intuitively check the potential for forming a relationship with the match through these results. The user can then proceed to follow-up conversations or recommendation activities based on the analysis results via phrases such as “Would you like to connect?”Accordingly, as illustrated in FIG. 5, the device of the present invention can provide the effect of personalizing the user experience and improving the matching persistence rate by not merely presenting matching results, but also visually guiding the potential for relationship development and personality compatibility with the matched partner.

[0069] In one embodiment, the device (100) may be configured to correct presentation bias, where user response varies depending on the placement position within the screen, even for the same matching candidate. For example, when selecting 20 top candidates and placing them in 6 slots of a user terminal, the past click / acceptance performance of each slot can be estimated in real time to continuously search for a certain proportion of positions so that a fixed position does not gain an excessive advantage. At the same time, diversity constraints can be applied to prevent excessively similar candidates from being clustered on the same screen, allowing the user to compare and evaluate the different strengths of multiple candidates. Since this screen-level learning operates as a presentation layer correction separate from the model's inherent matching score, it can provide the effect of increasing overall conversion (starting a conversation, meeting) through actual user response without compromising the fairness of the inherent matching logic.

[0070] In one embodiment, even if the same matching candidate is displayed at the top of the screen or in the first exposure area of ​​the user terminal, clicks or 'likes' may occur excessively due to visual attention regardless of content preference. Accordingly, the device (100) may include a presentation correction module to correct screen position bias. Specifically, the device (100) may collect screen position information where the matching candidate is displayed (e.g., top, middle, bottom slot position) and user response data for the candidate (clicks, conversation start, conversation duration, response rate, etc.). Subsequently, by comparing and analyzing the response data collected at different locations for the same candidate, the distortion of the click rate due to the position factor can be estimated, and the effect can be mathematically removed (debiased). For example, if the click rate is 30% when the same candidate is displayed at the top and 12% when displayed in the middle, the click rate increase effect (approx. 18%) due to the top display can be stored as a correction factor and reflected in subsequent evaluations. The data corrected in this way is reordered into “pure preference signals with location effects removed,” and the system can use location-corrected click-through rates or conversation retention rates as primary evaluation criteria when evaluating the performance of matching candidates in subsequent stages. That is, text and image signals of candidates that were simply “clicked because they were seen” are given lower weights as being temporary clicks, while signals verified by actual behavioral performance, such as when an actual conversation is maintained for a certain period of time after the click or when the mutual response rate is high, can be given higher weights.For example, if a user simply taps a profile picture displayed at the top but the conversation ends soon after, it is determined that the attractiveness factor of the image or text was temporarily overestimated, and the image / text-related weight of the candidate is adjusted; conversely, if a candidate is displayed relatively at the bottom and has a low exposure frequency but actually leads to a long-term conversation or mutual response, the weight of the candidate's behavioral signals (conversation maintenance, response stability, etc.) can be adjusted upward. In one embodiment, numerical examples such as 30%, 12%, etc. defined are merely examples and are not limited to these values.

[0071] Additionally, the device (100) may learn the accumulated click-through rates and conversation conversion rates for each location slot over a specific period to exploratorily optimize the screen layout itself. For example, if excessive click bias occurs in the top slot of the screen in a specific user group, the system may collect experimental data by randomly shuffling the candidate display order in some sessions and thereby improve the algorithm for estimating and correcting location effects. This feedback-based correction structure can effectively control visual bias occurring at the user interface (UI) stage without changing the matching logic of the model itself, thereby reducing the difference between actual likeability and matching quality. In one embodiment, the device (100) may monitor the distribution of recommendation frequency, matching success rate, and click-through rate relative to impressions for specific cohorts (e.g., same school, same region, specific age group, etc.) to quantitatively evaluate the fairness of the matching process, and may perform a warning or automatic correction if these indicators exceed a predefined deviation range. For example, if the ratio of the matching success rate to the recommendation exposure for specific cohorts A and B exceeds a preset threshold (e.g., 2 times), the device (100) may apply fairness constraints to slightly adjust the matching score for the cohort or to mitigate the distribution by limiting the recommendation frequency. At this time, fairness evaluation indicators may include disparity indicators (e.g., difference in success rates by group) and exposure balance indices (e.g., standard deviation of the average number of exposures for each group), and the device (100) may periodically calculate these indicators and provide them to the service operator for review at regular intervals. Consequently, the device (100) can produce sophisticated and fair matching results that reflect the potential for actual relationship formation by separating text and image signals into an “interest-generating stage” and behavioral signals into a “likability verification stage” through screen position bias correction.According to one embodiment, by utilizing multiple modal data such as text, images, and behavioral patterns to comprehensively identify user characteristics, it is possible to successfully perform matching between users with high potential intimacy rather than superficial matching. Furthermore, by utilizing user feedback (actual matching success / failure, whether to continue chatting, etc.) in reinforcement learning, a reinforcement learning-based self-evolving system can be implemented in which matching results change and improve in real-time according to user satisfaction. This improvement in matching accuracy is directly linked to user satisfaction and can bring about key commercial effects, such as increasing the service reuse rate and paid conversion rate.

[0073] Various embodiments of the present disclosure may be implemented as software comprising one or more instructions stored in a storage medium (e.g., memory) readable by a machine (e.g., a display device or a computer). For example, a processor (120) of the machine (e.g., processor (120)) may call at least one of the one or more instructions stored from the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' merely means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.

[0074] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created in a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0075] Although the present invention has been described with reference to the illustrated drawings, it is not limited by the disclosed embodiments and drawings, and those skilled in the art will understand that it may be implemented in modified forms without departing from the essential characteristics of the above description. Therefore, the disclosed methods should be considered in an illustrative rather than a restrictive sense. Even if the effects of the configuration according to the present invention are not explicitly described in the description of the embodiments, effects predictable by said configuration may also be recognized. The scope of the present invention is defined by the claims, not by the foregoing description, and all variations within the equivalent scope thereof should be interpreted as being included in the present invention. Explanation of the symbols

[0076] 100: Device 110: Receiver 120: Processor

Claims

Claim 1 A method for matching users based on artificial intelligence comprises: acquiring multiple modal data including profile text data, image data, and behavioral pattern data for multiple users; inputting the text data, image data, and behavioral pattern data into corresponding artificial intelligence models based on the multiple modal data to generate multiple modal feature vectors including text feature vectors, image feature vectors, and behavioral feature vectors; determining modal-specific weights for the multiple modal feature vectors based on at least one of user state and situational context, and mapping the multiple modal feature vectors with the applied modal-specific weights into a single integrated vector space through a multimodal fusion network; calculating similarity between users within the single integrated vector space to produce a matching score; and determining a matching result based on the matching score and providing it to a user terminal, and collecting user feedback data in real time and using the user feedback data for reinforcement learning to update the weights of the artificial intelligence model or the matching model in real time.The step of determining the weights for each modal includes assigning weights that gradually decrease in the order of the behavioral characteristic vector, the text characteristic vector, and the image characteristic vector to users whose accumulated amount of behavioral pattern data is greater than or equal to a preset threshold amount, and assigning weights that gradually decrease in the order of the image characteristic vector, the text characteristic vector, and the behavioral characteristic vector to new users whose accumulated amount of behavioral pattern data is less than the preset threshold amount, and if the situational context corresponds to a period where simultaneous access or fast response is restricted, lowering the weight corresponding to the response delay time among the behavioral pattern data and increasing the weight corresponding to the profile text data, and if the situational context corresponds to a late-night period or a period with frequent non-simultaneous access, increasing the weights corresponding to the profile text data and the image data for a certain period, and the step of updating the weights of the artificial intelligence model or matching model in real time calculates a first reward for whether mutual positive feedback between characteristic vectors is successful, a second reward for when the conversation duration is maintained for longer than a preset threshold time, a third reward for when the mutual response rate exceeds a threshold ratio, and a penalty for negative feedback based on the user feedback data, and assigns weights corresponding to the user vector pairs. A method for dynamically adjusting parameters.; Claim 2 A method according to claim 1, wherein the plurality of modal data further includes chat data, and the behavioral pattern data includes an indicator of user interaction derived by analyzing the chat data. Claim 3 In claim 1, the artificial intelligence model includes an artificial intelligence model for text processing, an artificial intelligence model for image processing, and an artificial intelligence model for behavioral pattern analysis, and the step of generating the multiple modal feature vector comprises: the artificial intelligence model for text processing including a large-scale language embedding model specialized for Korean to generate semantic embedding vectors from user profile sentences to generate the text feature vectors; and the artificial intelligence model for image processing including a vision-language integration model to generate image embedding vectors representing visual tendencies from user profile images to generate the image feature vectors. Claim 4 In claim 3, the step of generating the multiple modal characteristic vectors is a method in which the artificial intelligence model for behavioral pattern analysis receives the user's in-app activity logs, conversation frequency, response latency, and movement pattern data and generates the behavioral characteristic vectors. Claim 5 In claim 1, the multimodal fusion network includes an attention mechanism or a transformer-based fusion module, and the step of mapping to a single integrated vector space comprises generating an optimal integrated vector by assigning differential weights to feature vectors among the multiple modal feature vectors that contribute significantly to matching accuracy, and mapping the result to the single integrated vector space. Claim 6 The method according to claim 1, wherein the step of calculating the similarity to produce a matching score comprises: a step of calculating the similarity using a hybrid calculation method combining cosine similarity and weighted Euclidean distance; and further comprises a step of modeling user relationships within the single integrated vector space as a graph neural network and predicting potential favorability and relationship development possibilities between the users. Claim 7 An AI-based user matching device comprises: a receiver that acquires multiple modal data including profile text data, image data, and behavioral pattern data for multiple users; and a processor that inputs the text data, the image data, and the behavioral pattern data into corresponding AI models based on the multiple modal data to generate multiple modal feature vectors including text feature vectors, image feature vectors, and behavioral feature vectors, determines modal-specific weights for the multiple modal feature vectors based on at least one of user state and situation context, maps the multiple modal feature vectors with the applied modal-specific weights into a single integrated vector space through a multimodal fusion network, calculates similarity between users within the single integrated vector space to produce a matching score, determines a matching result based on the matching score and provides it to a user terminal, and collects user feedback data in real time and updates the weights of the AI ​​model or the matching model in real time as the user feedback data is used for reinforcement learning.A device comprising: a processor that assigns gradually decreasing weights in the order of the behavioral characteristic vector, the text characteristic vector, and the image characteristic vector to users whose accumulated amount of behavioral pattern data is greater than or equal to a preset threshold amount; assigns gradually decreasing weights in the order of the image characteristic vector, the text characteristic vector, and the behavioral characteristic vector to new users whose accumulated amount of behavioral pattern data is less than the preset threshold amount; if the situation context corresponds to a period where simultaneous access or fast response is restricted, lowers the weight corresponding to the response delay time among the behavioral pattern data and raises the weight corresponding to the profile text data; if the situation context corresponds to a late-night period or a period with frequent non-simultaneous access, raises the weights corresponding to the profile text data and the image data for a certain period; and, based on the user feedback data, calculates a first reward for whether mutual positive feedback between characteristic vectors is achieved, a second reward for when the conversation duration is maintained for longer than a preset threshold time, a third reward for when the mutual response rate exceeds a threshold ratio, and a penalty for negative feedback, thereby dynamically adjusting weight parameters corresponding to user vector pairs.

Citation Information

Patent Citations

  • Multimedia recommendation method and apparatus for relrecting modality characteristics and modeling user interest on target items

    KR1020240048815A

  • Manager Matching Method Using Marriage Information Matching Online Service Platform

    KR102092906B1

  • Method, apparatus, and system for providing an ai-based business and manufacturer matching platform service

    KR102809718B1