System and method for selecting relevant multimedia content
The system addresses the challenge of providing relevant multimedia content by combining assessments of invariant and variable human traits to deliver personalized content based on real-time data, enhancing user engagement and campaign effectiveness.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2026-04-02
AI Technical Summary
Existing content distribution systems fail to provide relevant multimedia content to individuals by not considering invariant and variable human characteristics, such as appearance and emotions, leading to ineffective and potentially off-putting information presentation.
A system and method that combines the assessment of invariant traits (appearance) and variable traits (emotions) to determine a person's psychotype, using neural networks to provide personalized multimedia content based on real-time data from sensors and video recordings without conscious user interaction.
Enables objective and reliable selection of multimedia content, improving user perception and engagement by providing relevant stimuli based on accurate assessments of a person's characteristics, reducing emotional distortion and enhancing the effectiveness of advertising campaigns.
Smart Images

Figure 00000032_0000 
Figure 00000033_0000 
Figure 00000034_0000
Abstract
Description
[0001] SYSTEM AND METHOD FOR SELECTING RELEVANT MULTIMEDIA CONTENT
[0002] AREA OF TECHNOLOGY
[0003] The proposed invention relates to the field of multimedia content management, and in specific embodiments to technologies and mechanisms for selecting the most significant multimedia content for a person, ensuring the identification of the expected characteristics of a person, on the basis of which the most relevant multimedia content for a given person is selected.
[0004] LEVEL OF TECHNOLOGY
[0005] As information systems evolve, consumer perceptions of information are changing, with video content attracting the greatest audience attention. Therefore, information distributors and / or providers, whether advertising, educational, or other content, typically need to present content onscreen to attract audience attention and achieve desired results, such as sales promotion or knowledge acquisition. The connection between information products and consumers can be improved through intelligent distribution of personalized content, ensuring timely and relevant information delivery to potential consumers by selecting the time and place to display content to potential audience members (target audiences) based on various factors, such as demographics, purchase history, learning dynamics, observed viewing behavior, and others.
[0006] Today, the state of the art provides various methods and systems that allow content distributors to target specific audiences that are most receptive to the information provided.
[0007] A personalized advertising promotion system based on big data and artificial intelligence is known (CN 116629942 A), comprising: a data collection module used to obtain data on supermarket visitors, sales data and supermarket opening hours within a standard time period T, and to send data on the flow of supermarket visitors to a data processing module; a data processing module used to correspondingly process and analyze the received data on the flow of supermarket visitors and sales data; a data comparison module used to sort and calculate weight types corresponding to time periods based on the visitor flow coefficient and the area ratio of the corresponding time periods and to output the weight types to the advertising management module;An advertising management module used to transmit the appropriate types of advertisements to the advertising playback module at the appropriate time periods and manage the advertisements to be played. A drawback of the existing system is that it uses supermarket traffic data in its analysis without analyzing individual shopper characteristics, including their personality types and moods. In this case, the user may not necessarily be presented with relevant information, as it is impossible to control factors in advertising perception, such as, for example, the variability of information presentation. Consequently, the information displayed may be off-putting to the user.
[0008] Furthermore, a system and method for transmitting personalized content in communication devices (W02009101227A1) is known in the prior art. The content is delivered via one or more mobile networks, allowing users to receive personalized advertising selected from a distribution center based on each user's preferences, with said information collected and processed by artificial intelligence or a statistical unit associated with the communication display device.
[0009] The disadvantages of this well-known technology include the lack of a developed sensor system, the lack of the ability to interact with a person, and the limitation of interaction with a person only via a mobile network.
[0010] Another analogue of the proposed invention is a method and system for managing advertising placement based on artificial intelligence (CN117078312A). The invention describes algorithms for processing data characterizing human behavior, but does not address the issue of presenting visual information to people based on their behavior and subsequently analyzing their behavior based on the presented information.
[0011] This known solution concerns only the analysis of user behavior, but not the system for presenting informational, advertising or educational materials in general, and therefore does not have all the advantages of the claimed invention.
[0012] The closest analogue of the proposed invention is the solution according to RU2672171C1, which discloses a method for processing information for preparing recommendations based on a computerized assessment of user abilities, in which access is provided using authentication means to the solution of at least one computerized task to the user's device, the user's emotions are recognized in real time by means of one video camera, with the help of which a set of emotions is obtained and artificial neural networks are used to recognize the received emotions, the data is transmitted to the data processing device of the system to obtain the parameters of the user's actions during the said decision, at least one user ability is determined based on the data obtained after the user's decision and the recognized user emotions, in the data processing device the data is converted into numerical values, and the recognized emotions are converted into a scale,which are used to calculate the user's ability score, and based on the results obtained, the system generates recommendations for the user to make decisions.
[0013] The disadvantages of the closest analogue include the requirement for user authentication, after which the user is engaged in conscious interaction with the system to assess their abilities. Consequently, the user is required to perform a number of actions that influence the subsequent assessment and its results. In particular, conscious interaction with the system can change a person's characteristics, which are then assessed by the system. In other words, the neural network takes into account the user's characteristics, but it ignores the fact that some of the user's demonstrated characteristics are irrelevant to the assessment process.
[0014] In the proposed invention, the user's participation is unconscious; the user's evaluation occurs without the user's conscious interaction with the system for authentication in the system and for other purposes.
[0015] The following sources are known from non-patent literature.
[0016] The Microsoft blog published an article titled "Psychological Portrait Using a Neural Network and a Regular Camera" (https: / / habr.com / ru / companies / microsoft / articles / 348966Z), which describes approaches aimed at finding people matching a user's interests and personality, whether friends or significant others, based on social media data. In the article, the camera's role is limited to recognizing a person's appearance for the purpose of searching for information about them on the internet and social media, while the psychological type is determined based on the semantics of the user's texts left on social media. This approach does not necessarily yield reliable results, as the user's style of expression is debatable as an indicator.
[0017] The publication "Extrovert or Introvert? Neural Networks Will Help Detect Personality Psychotype" (https: / / www.susu.ru / ru / news / 2018 / 01 / 12 / ekstravert-ili-introvert-neyroseti-pomogut-vychislit-psihotip-lichnosti) mentions that scientists at South Ural State University, while implementing a project on artificial neural networks, "Diagnostics of Human Psychological Type According to Jung," developed and implemented methods and algorithms for automatically determining a person's psychotype based on the analysis of a video recording of the subject's psychological testing. In this case, the psychotype is determined based on the analysis of a video recording of the subject's psychological testing. This is a rather lengthy and labor-intensive process that does not provide for prompt interaction with the person.
[0018] A source (https: / / iq-media.ru archive / 366225961.html) is also known, describing how a cascade neural network was trained to identify the Big Five personality traits (BIG5) from a facial photograph. The Big Five is a model for assessing individual differences in human personality based on five dimensions. These include extroversion, openness, agreeableness, conscientiousness, and neuroticism. These characteristics may or may not be present in any personality, to varying degrees. The cascade neural network was provided with a pre-labeled dataset. It consisted of test results and 31,367 photographs from 12,500 volunteers. After training, the algorithm was able to predict five common personality traits. The neural network was best at identifying conscientiousness and conscientiousness. Initially, the neural network was trained to distinguish between the faces of different people, but to consistently recognize the face of the same person.The algorithm was then trained to decompose each image into 128 invariant features—regularly recurring individual characteristics. Within the model, each invariant was represented as a vector in a 128-dimensional space. The full study is also published on the website of the journal Scientific Reports (https: / / www.nature.com / articles / s41598-020-65358-6).
[0019] This algorithm recognizes only invariant human characteristics, namely facial features, from a person's photograph, and on this basis determines the Big Five personality traits from a facial photograph.
[0020] The publication "Behavioral Analytics in Video Surveillance" (https: / / videomir.pro / stati / povedencheskaya-analitika-v-videonablyudenii / ) examines the possibility of using neural networks to analyze human behavior and create behavioral detectors for various types of people. This technology could be used in security, for example, to extract from a general scene an image that depicts a particular event and the people present; to detect various special objects left in public places with malicious intent; to detect the presence of people in places where their safety and health are threatened by various mechanical or weather events; and to monitor the implementation of safety measures and compliance with technological processes.
[0021] This technology is limited in that it only allows one to identify certain behavioral characteristics of a person and classify them as undesirable without identifying any other characteristics of that person, including their psychotype.
[0022] The publication htps: / / tass.ru / obschestvo / 6328823 provides examples of how cameras in supermarkets recognize people's emotions and explains what emotions the neural network detects and how this works. American psychologist Paul Ekman identifies seven basic human emotions: anger, disgust, fear, happiness, sadness, surprise, and neutral. The algorithm identifies faces in a crowd, captures "points" on them—that is, areas where certain facial expressions are evident—and combines these features. It then goes to a photo database to compare this unidentified image with images in which the emotion has already been identified, compares the image, and determines the emotion. However, the algorithm is less effective at identifying emotions than invariant features such as age, gender, and others. Thus, the state of the art currently includes specific algorithms that recognize:
[0023] - human emotions, that is, variable features;
[0024] - features of a person’s appearance, including both the face and the person’s figure as a whole, that is, invariant characteristics;
[0025] - a person's psychotype.
[0026] One of the challenges of providing users with relevant informational content is the need to process the enormous amount of data collected about the user and their behavior. In reality, this still doesn't provide an understanding of what interests a person at a given moment; even users themselves don't always understand what they like and why. This is why the proposed invention utilizes unconscious user participation; the user is initially located in a public space and has no intention of interacting with the system.
[0027] Another challenge is improving user perception of displayed content. This issue has generally received insufficient attention. As is clear from the aforementioned prior art sources, only one of the following is assessed: either the user's emotional response, or a psychotype or similar personality model, such as the Big Five or Paul Ekman's traits; or invariant human traits, with these assessments being conducted for different purposes. Some sources assess invariant human traits (appearance) alongside variable traits (emotions), and practice decision-making based on a combination of these assessments.However, these sources fail to address the issue of providing relevant content to a person; nor do they mention the effectiveness and speed of such combined human assessments for providing operational information about a person and subsequent interaction with the person by showing them hypothetically relevant content based on the information received about the person. At the same time, they mention that some algorithms recognize certain emotions better than others, or even confuse emotions. For example, neural networks are best at recognizing when people are happy. Neutral emotions are often confused with negative emotions, while this is a significant omission when deciding which content to show to a person.
[0028] In general, to solve the existing problems of providing users with relevant content, it is necessary to ensure a more precise selection of relevant content (stimuli) by taking into account additional human characteristics, such as emotions, mood, and so on. Furthermore, it is useful to assess human characteristics without their involvement, that is, without requiring them to undergo any testing or notifying them that they are being monitored by a neural network. By taking into account additional human characteristics without their awareness that additional characteristics are being assessed, the person subsequently experiences less psychological barriers and biases toward the information being presented, such as the advertised object or educational material. This also addresses the issue of improving the user's perception of the advertised object or educational material.Thanks to the presented stimuli, the user can be engaged in a dialogue and, over the course of the dialogue, improve their attitude toward the advertised item. In this way, the user expresses their sincere and frank emotions, or emotions that are likely more sincere than in the case of direct contact with the algorithm. Thus, emotional distortion is reduced and the likelihood of obtaining more reliable characteristics of the person is high, thereby ensuring the formation of specific, relevant informational content for the user, both verbal and nonverbal.
[0029] The proposed invention makes it possible to overcome at least some of the disadvantages of technical solutions known from the prior art.
[0030] As follows from the following description of the invention, the technology according to the proposed invention ensures the determination of primary and secondary characteristics of a person.
[0031] The following types of neural networks can be implemented: multilayer perceptron and backpropagation algorithm.
[0032] ESSENCE OF THE INVENTION
[0033] The solution to these and possibly other problems is a combined assessment of invariant and variant traits, with a separate assessment of a person's psychotype or similar personality model, such as the Big Five or Paul Ekman's traits. In this patent application, the combination of invariant traits such as appearance and variable traits such as facial expressions is referred to as "primary characteristics." In this patent application, a psychotype or similar personality model, such as the Big Five or Paul Ekman's traits, is referred to as "secondary characteristics."
[0034] The technical result is the provision of objective and reliable information about the target audience for the display of advertising, informational, or educational content, enabling an objective assessment of the effectiveness and efficiency of advertising campaigns, the dynamic selection and playback of advertising and informational advertisements in accordance with the characteristics of the potential consumer, and an increase in the relevance of visual stimuli presented to a person, based on the external visible characteristics of the person present in a public place.
[0035] The technical result is achieved thanks to the proposed method and system.
[0036] According to a first aspect of the present invention, a method is proposed for providing relevant content based on characteristics of a person, which includes: collecting data about at least one person using at least one sensor device and at least one video recording device; transmitting data about said at least one person to a server; interpreting data about said person on the server by first obtaining primary characteristics and secondary characteristics of said at least one person; making a decision to demonstrate a first multimedia content to said at least one person based on the first received primary characteristics and secondary characteristics of said at least one person; performing a first multimedia operation, including demonstrating the first multimedia content to said at least one person using an information output device;second obtaining primary characteristics and secondary characteristics of said at least one person; making a decision to demonstrate a second multimedia content to said at least one person based on the second obtained primary characteristics and secondary characteristics of said at least one person.
[0037] According to a non-limiting embodiment of the invention, the proposed method involves assigning a conditional identifier to at least one person who has fallen into the field of observation of sensor devices and video recording devices.
[0038] According to a non-limiting embodiment of the invention, in the proposed method, people who fall within the observation field of sensor devices and video recording devices are combined into a group with subsequent obtaining of variable and / or invariant features of at least one person from the group.
[0039] According to a non-limiting embodiment of the invention, the proposed method receives feedback from said at least one person using sensor devices and video recording devices in the form of a reaction to the first relevant multimedia content.
[0040] According to a non-limiting embodiment of the invention, the proposed method additionally uses feedback from a person to use the obtained human characteristics in further training the data analysis algorithm.
[0041] According to a non-limiting embodiment of the invention, in the proposed method, the data analysis algorithms are neural networks.
[0042] According to a non-limiting embodiment of the invention, the content is displayed by selecting from an existing database or generating it using neural networks.
[0043] According to a non-limiting embodiment of the invention, in the proposed method, at least one sensor device is a microphone, and at least one video recording device is a video camera.
[0044] According to a non-limiting embodiment of the invention, the sensor devices in the proposed method may comprise a sensor system. According to a non-limiting embodiment of the invention, the video recording devices in the proposed method may comprise a video camera system.
[0045] According to a non-limiting embodiment of the invention, in the proposed method, the video recording device is calibrated to ensure the possibility of ensuring: capturing an image of a person; sufficient resolution and sensitivity for collecting data about a person.
[0046] According to a non-limiting embodiment of the invention, in the proposed method, the human characteristics may be variable, invariant, and invariant characteristics.
[0047] According to a non-limiting embodiment of the invention, in the proposed method, secondary characteristics of a person are calculated based on primary characteristics of a person.
[0048] According to a non-limiting embodiment of the invention, in the proposed method, the primary characteristics are at least
[0049] - facial expressions;
[0050] - the relative position of body parts and parameters of human movement;
[0051] - color characteristics, including clothing color, hair color, eye color;
[0052] - gender, age, anthropometric parameters; and secondary characteristics represent at least the person’s psychotype.
[0053] According to a non-limiting embodiment of the invention, the proposed method additionally comprises:
[0054] - the third obtaining of primary characteristics and secondary characteristics of the said at least one person;
[0055] - making a decision on demonstrating the third multimedia content to the specified at least one person based on the second received primary characteristics and the secondary characteristics of the specified at least one person.
[0056] According to a second aspect of the present invention, there is provided a system comprising
[0057] - server;
[0058] - sensor devices and video recording devices, designed with the ability to activate and monitor a person and connected to the server via a communication channel;
[0059] - information output devices designed to enable the display of multimedia content to a person;
[0060] - computing modules designed with the ability to execute a data processing algorithm and receive data from sensor devices and / or video recording devices;
[0061] - a database intended for storing data and connected to computing modules; wherein the computing modules include a module for identifying the primary characteristics of a person and a data analysis module; wherein the module for identifying the primary characteristics of a person is configured to receive data from sensor devices and / or video recording devices and, based on this, to calculate the primary characteristics of a person; the data analysis module is configured to: calculate the secondary characteristics of a person based on the primary characteristics of a person; generate recommendations for relevant multimedia content based on the calculated primary and secondary characteristics of a person; a data output module configured to demonstrate relevant multimedia content to a person using an information output device.
[0062] According to a non-limiting embodiment of the system, the communication channel through which the sensor devices and video recording devices are connected is wired or wireless.
[0063] According to a non-limiting embodiment of the system, at least one sensor device is a microphone, and at least one video recording device is a video camera.
[0064] According to a non-limiting embodiment of the system, the sensor devices may be a system of sensor devices.
[0065] According to a non-limiting embodiment of the system, the video recording devices may be a system of video cameras.
[0066] According to a non-limiting embodiment of the system, the video recording device is calibrated to ensure the possibility of providing:
[0067] - capturing a human image;
[0068] - sufficiency of resolution and sensitivity for collecting data about a person.
[0069] According to a non-limiting embodiment of the invention, the primary and secondary characteristics are calculated in real time. This enables the timely calculation of a person's primary and secondary characteristics, the determination of an optimal visual stimulus corresponding to these characteristics, and its presentation to the person while the person is within the camera's field of view and / or microphone's range and their attention is focused on the information output device. The information may be text and / or an image. The information output device may be a television, a projector combined with a surface for displaying images and other information, a laptop screen or computer monitor, a tablet screen, a mobile device screen, or any other surface capable of displaying information, without limitation.And also represent LED, OLED or other screens, which is obvious to a specialist in this field of technology.
[0070] Preferably, the proposed system comprises audio input and output devices configured to output and receive audio content.
[0071] In the context of this application, the relevant multimedia content is also referred to as the "stimulus".
[0072] In the context of this description, unless specifically stated otherwise, the words "first," "second," "third," etc. are used as adjectives solely to distinguish the nouns to which they refer from one another, and not for the purpose of describing any specific relationship between these nouns. For example, it should be kept in mind that the use of the terms "first server" and "third server" does not imply any order, classification, chronology, hierarchy, or ranking (for example) of servers / between servers, nor does their use (in itself) imply that some "second server" necessarily exists in a given situation. Further, as noted herein in other contexts, reference to a "first" element and a "second" element does not exclude the possibility that they are one and the same actual, real element.For example, in some cases, the “first” server and the “second” server may be the same software and / or hardware, while in other cases they may be different software and / or hardware.
[0073] BRIEF DESCRIPTION OF DRAWINGS
[0074] The accompanying drawings, which are included in this description and form a part thereof, illustrate one or more embodiments of the technology objects together with a detailed description, serve to explain the principles and embodiments of the technology.
[0075] Fig. 1 shows a schematic representation of a system according to the invention.
[0076] Fig. 2 schematically shows a block diagram of the method according to the invention.
[0077] Fig. 3 is a schematic diagram of the primary characteristics used in the invention.
[0078] Fig. 4 shows a schematic diagram of the secondary characteristics used in the invention. The class of secondary characteristics is emotions. Fig. 5 shows a schematic diagram of the secondary characteristics used in the invention. The class of secondary characteristics is psychotype.
[0079] Fig. 6 shows a table of primary characteristics for determining emotions based on facial expressions, used in the invention.
[0080] Fig. 7 is a schematic representation of a class diagram including visual stimuli, input data, human observation base, and operator participation.
[0081] Fig. 8 is a schematic diagram of the system components showing what the system consists of.
[0082] Fig. 9 schematically shows a block diagram of the additional training used in the invention.
[0083] DETAILED DESCRIPTION OF THE INVENTION WITH REFERENCES TO DRAWINGS
[0084] Examples of the present technology are described herein in the context of systems, methods, and computer program products to provide relevant information to users. Those skilled in the art will appreciate that the following description is illustrative only and is not intended to be limiting. Other embodiments of the invention will be apparent to those skilled in the art having the benefit of this description. Embodiments of the aspects illustrated in the accompanying drawings will now be described in more detail. Throughout the drawings and in the following description, the same elements will be numbered identically whenever possible.
[0085] Fig. 2 shows a block diagram of the implementation of the method for providing relevant content.
[0086] In the first stage, data is collected about at least one person using at least one sensor device and at least one video recording device. The sensor devices and video recording devices may be represented by a system of microphones and cameras, respectively. It will be apparent to those skilled in the art that sensor devices such as motion sensors or presence sensors may also be used. The video cameras used may be connected to the server via a wired or wireless communication channel. Several different connection options are possible. In one embodiment, the video cameras are directly connected to the server. In another embodiment, communication between the video cameras and the server is accomplished via an internal network built using network equipment.In another embodiment of the invention, communication between the video cameras and the server is accomplished via public networks (including the internet), with the video cameras having specific network addresses and authentication parameters for access. When connecting via public networks, the server can be located remotely from the video cameras. The choice of connection method for a specific product is determined by the design features of the video cameras used. One possible solution is to install a laptop with a built-in video camera and microphone, which are activated by the presence of people in the room.The term "module" in this context means a physical device, apparatus, or a plurality of modules implemented using hardware, such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA), or a combination of hardware and software, such as a microprocessor system and a set of instructions implementing the module's functionality, which (when executed) transform the microprocessor system into a special-purpose device. A module may also be implemented as a combination of two modules, with certain functions supported only by hardware and other functions supported by a combination of hardware and software. In some specific embodiments of the technology, at least some of the modules (and in some cases, all) may be used on a general-purpose computer processor.Accordingly, each module may be implemented in a variety of different configurations and is not limited to the particular embodiment given herein as an example.
[0087] The server includes a microprocessor that executes software instructions. The microprocessor is connected to a storage medium containing databases with information about the consumer profile and information content, as well as software instructions that the microprocessor executes. The devices (modules) and / or their constituent parts described within the framework of the present invention can be connected to each other via wired and / or wireless communication means / methods, as well as via various types of connections, including detachable or permanent (e.g., via terminals, contacts, adapters, soldering, mechanical connections, etc.). Communication means can be implemented, for example, via local area networks (LAN), a USB interface, an RS-232 interface, Bluetooth, Wi-Fi interfaces, the Internet, and mobile cellular communications (GSM).Data transfer between modules and devices of this technical solution can be carried out via HTTP, HTTPS, FTP, TCP / IP, etc. protocols, including IEEE 802.15.4 and ZigBee protocols, including APS and NWK using lower-level services, the MAC medium access control layer and the PHY physical layer, and others.
[0088] A server can be either a single physical server or a virtual server, which is a group (cluster) of physical servers connected via network links. Such a group of servers can also be a cloud solution. In this case, the server can utilize both central processing units (CPUs) and graphics processing units (GPUs) for computing.
[0089] According to the second aspect of the invention, a system 1 is proposed (see Fig. 1) comprising a server 2; at least one sensor device 3 and at least one video recording device 4. According to a non-limiting embodiment of the invention, the sensor device 3 is a microphone, and the video recording device 4 is a video camera. The microphone 3 and the video camera 4 are configured to activate and monitor a person 5 and are connected to the server 2 via a communication channel 6. The information output devices 7 are configured to demonstrate multimedia content 8 to a person 5 and are, according to a non-limiting embodiment, various monitors or screens, as well as any surface onto which an image is projected. Liquid crystal monitors (LCD), plasma monitors (PDP), OLED monitors, QLED monitors and other monitors known to a person skilled in the art can be used.
[0090] System 1 comprises computing modules capable of executing a data processing algorithm and receiving data from microphone 3 and video camera 4. The computing modules include module 9 for identifying primary human characteristics 5 and module 10 for data analysis. System 1 comprises database 11 intended for storing data and connected to the computing modules. Module 9 for identifying primary human characteristics is capable of receiving data from microphone 3 and video camera 4 and, based on this, calculating primary human characteristics. Data analysis module 10 is capable of calculating secondary human characteristics based on the calculated primary human characteristics. System 1 also comprises data output module 13.
[0091] According to a first aspect of the invention, a method for providing relevant content based on characteristics of a person is proposed, comprising the following steps.
[0092] Step 1. Collecting data about at least one person 5 using at least one sensor device 3 and at least one video recording device 4; the data about at least one person 5 may be an image from video cameras 4, a recording from microphones 3, static data of the person 5, a continuous video stream of the person's speech 5.
[0093] The input data of this invention is sensor data, namely audio recording from microphones, images from cameras (digital matrix or 3D image), and also works on image data using a machine learning (ML) model, taking static data or continuous video stream as input and outputting a list of detection results.
[0094] Stage 2. Transfer of data about at least one specified person 5 to server 2. Data transfer is carried out via communication channel 6.
[0095] Step 3. Interpreting data about the specified individual on server 2, first obtaining the primary characteristics and secondary characteristics of at least one specified individual. The proposed invention utilizes methods for calculating the primary characteristics of an individual and the secondary characteristics of an individual.
[0096] Methods for calculating secondary characteristics are primarily methods for classifying multivariate data. Various methods are used for these tasks (decision trees, Bayesian classifiers, support vector machines, multilayer neural networks, etc.). The algorithm's input is a vector of values of a person's primary characteristics, calculated in previous steps. The output is a classification result, which is the value of the resulting secondary characteristic or a vector of probabilities for different characteristic values. The components of this vector can be either numeric or categorical variables.
[0097] During data processing, in addition to directly calculating primary and secondary characteristics, a number of related tasks are solved. These include image segmentation and human detection, noise and object removal, keypoint detection, and more.
[0098] According to Fig. 6, there is a set of features that are part of the secondary feature vector and allow for the detection and extraction of human mood features from an image using given emotion characteristics that are determined by facial features such as the eyes, mouth, eyebrows, cheeks, nose, as well as facial expressions and other details. According to Paul Ekman (Ekman, Paul. Emotions Revealed: Recognizing Faces and Feelings to Improve Communication and Emotional Life I Transl. from English: V. Kuzin. - St. Petersburg: Piter, 2010. - 336 p. - ISBN 978-5-498-07705-5) there are seven basic emotions: happiness, surprise, fear, anger, sadness, disgust, contempt.
[0099] This task operates on image data using a machine learning (ML) model, accepting static data or a continuous video stream as input and outputting a list of detection results. The result is a human emotion detection. The algorithm analyzed training samples, specifically a set of photographs, and identified the distinctive features of each emotion. Figure 6 illustrates emotion detection based on facial expressions in tabular form. The program step-by-step determines which emotion is represented by a feature associated with a facial expression (eyes, mouth, eyebrows, cheeks, nose, etc.).
[0100] A person is considered to be feeling sad when the upper eyelids are slightly drooping, the lower ones are relaxed, there is a distracted gaze directed downwards at one point, the eyes themselves are half-closed, while the corners of the mouth are slightly lowered, and the lips are relaxed, the mouth is closed, the lower lip can be slightly pushed forward, and the eyebrows are raised.
[0101] The emotion of happiness can be seen by the small wrinkles at the corners of the eyes, including the muscles around the eyes. The muscles at the corners of the eyes are tense, as if the person is squinting slightly. The lips are relaxed, the smile is symmetrical, the eyebrows are relaxed, and the cheeks are raised. When a person smiles sincerely, the outer part of the muscle is engaged; a true smile involves all the facial muscles.
[0102] When a person is angry, their eyes are open as wide as their lowered eyebrows will allow, their gaze is downcast, their mouth is closed, their lips are slightly pursed and narrowed, the corners of their lips are turned down, their lips are tense, and their teeth are clenched. Typically, their eyebrows are lowered, tense, and drawn together, with a deep crease at the junction of their eyebrows above the bridge of their nose. Their cheeks are tense, and their face is flushed and slightly tilted forward.
[0103] The emotion of fear is usually visible by the raised upper eyelids while the lower eyelids are tense, the eyes are wide open, the mouth is slightly open, and the lips are tense and slightly stretched. The eyebrows are raised, narrowed, and elongated, while the facial muscles are generally tense.
[0104] All facial muscles are involved in an expression of disgust, and the face wrinkles: the eyes are narrowed sharply, the upper lip is raised and tense, and the corners of the lips are drawn down. The mouth is slightly open, exposing the upper teeth due to the raised upper lip. Meanwhile, the eyebrows are lowered and drawn together, a crease forms between the eyebrows, the cheeks are raised, and the nose and bridge of the nose are wrinkled.
[0105] Surprise can be seen in the wide-open eyes, slightly open mouth, relaxed lips, strongly lowered chin, raised eyebrows, and relaxed facial muscles.
[0106] Primary human characteristics, as mentioned earlier, include invariant features such as clothing color, hair color, eye color, gender, age, and anthropometric parameters, as well as variable features such as facial expressions. Primary characteristics are described by sets of numerical and categorical variables (see Fig. 3). Data classification algorithms based on convolutional neural networks operate in the primary characteristic recognition units.
[0107] Those skilled in the art will recognize a variety of algorithms for real-time object recognition in images. Therefore, they can select the appropriate one for the specific conditions under which the method according to the invention will be implemented. In particular, the Viola Jones algorithm is fundamental for real-time object recognition in images and enables the recognition of objects at low angles under various lighting conditions. It is also one of the best in terms of recognition efficiency and speed. Furthermore, those skilled in the art will appreciate the additional use of boosting algorithms to ensure the consistent composition of machine learning algorithms, so that each subsequent algorithm seeks to compensate for the shortcomings of all previous algorithms.
[0108] The data classification algorithms used here are pre-trained on training sets. Both ready-made labeled datasets and datasets specifically prepared for this work are used as training sets. These datasets may contain data on the correspondence of images and videos to skeletonized or mesh images, various categories of human gestures and poses, and various types of facial expressions, clothing, and human behavior.
[0109] In addition to data classification algorithms, these blocks employ a number of auxiliary algorithms for data processing and preparation, including noise removal, highlighting significant image fragments, etc.
[0110] Module 9 of identification of primary human characteristics allows to detect:
[0111] - facial features, including without limitation facial expressions;
[0112] - the relative position of body parts and the parameters of human movement, including, without limitation, gestures;
[0113] - color characteristics of a person, including, without limitation, color of clothing, hair color, eye color;
[0114] - gender, age, anthropometric parameters without restrictions.
[0115] For the software implementation of the module for identifying primary and secondary human characteristics, ready-made library solutions are used, such as OpenCV, TensorFlow, PyTorch, and others, which a specialist in this field of technology can select independently.
[0116] The second task is to recognize the secondary characteristics of a person, the psychotype (see Fig. 4 and Fig. 5).
[0117] The characteristics or traits used to determine a person's psychotype include mood, gait, clothing style and color, behavior, and so on. Our analysis revealed that different psychotypes respond best to different visual stimuli. Data analysis module 10 allows you to identify a person's psychotype, their psychological profile, using preset characteristics of various psychotypes, which will be described below.
[0118] 1) According to Hippocrates: choleric, i.e. emotional, active and courageous, leader and initiator; sanguine, i.e. calm and balanced, optimist and “life of the party”; phlegmatic, i.e. calm and reserved, sedentary; melancholic, i.e. pessimist, prone to despondency, creative person;
[0119] 2) According to Jung: introvert (keeps emotions under control, closed, perfectionist), extrovert (stage people, speakers, optimists);
[0120] 3) According to Leonhard: hyperthymic (active, sociable), dysthymic (isolation, poor contact with people), stuck (stuck on negative emotions, vindictive), anxious (timid and withdrawn, friendly), demonstrative (center of attention, intriguer), excitable (prone to conflict, irascibility), pedantic (neat and demanding), cycloid (quickly changing mood), emotive (responsive, understanding, non-conflict), exalted (sociable, gets attached to people)
[0121] 4) According to Fromm: receptive (dependent on people's opinions), exploitative (achieve everything), accumulating (material goods are important), market type (for them, personality is a commodity), productive (independent, honest, calm); 5) According to Rotter: external (victim of circumstances, adapts poorly), internal (everything depends on him, soberly assess reality);
[0122] 6) According to Lichko and Gannushkin: hysteroids (must be the center of attention), paranoids (goes to the goal over bones), epileptoids (gloomy irritability, hostility towards others), emotives (modest, independent), hyperthymics (always on the move, constantly doing something, do not like to be in the shadows, love adventures), schizoids (seeks protection, lives in their own world, lonely).
[0123] The psychotype, which in the context of this application is also referred to as “secondary characteristics”, can be calculated on the basis of primary characteristics.
[0124] Also, to accurately determine a person's age, the characteristics of each class are taken into account, which can help the algorithm:
[0125] - in children this is short stature, a round face, children's clothing, specific attributes such as toys in hands or a pacifier;
[0126] - in teenagers this may be short stature, school uniform, thoughtless actions, frequent mood swings, loud expression of emotions, causing self-confidence, excessive gesticulation, experimentation with appearance, a desire to be part of a group;
[0127] - as for adults, these characteristics are body type, conscious behavior, emotional security, and perception of humor;
[0128] - for pensioners, these parameters may be: wrinkles, gray hair, short stature, slow gait, and the presence of a cane or walker.
[0129] A significant number of methods for calculating primary human characteristics are based on the use of convolutional neural networks. The input to such a neural network is an image represented by a set of pixels (or, in the case of video analysis, a set of frames, each representing a set of pixels). The output is a classification result, which represents the value of the obtained primary characteristic or a vector of probabilities for different characteristic values. The components of this vector can be either numerical or categorical variables.
[0130] Step 4. Making a decision on demonstrating the first multimedia content to at least one specified person based on the first received primary characteristics and secondary characteristics of at least one specified person 5.
[0131] Step 5: Performing a first multimedia operation, including demonstrating to said at least one person the first multimedia content using the information output device.
[0132] Visual stimuli. The invention is described below with reference to Fig. 7. Visual stimuli can be represented by either images or videos, and can also engage in dialogue with a person, provide advice on various issues, and communicate. Video clips are assembled from fragments (divided into groups and then selected from the desired group individually for each person). Visual stimuli are formed using video archive data containing templates and components of visual stimuli. The components of visual stimuli include, in particular, skeletons of VR images and textures. Both videos from the video archive and generated VR models are used.
[0133] This may be not only an image of a person, but also an animal or a fictional character. Visual stimuli can be displayed based on the system's focus of attention on a specific person and can alternate between different people, taking into account the system's shifting focus.
[0134] They can also be expressed in both text and speech. The semantics of messages displayed in conjunction with visual stimuli are selected based on the primary and secondary human characteristics calculated in the previous stages in such a way as to enhance the effect of visual stimuli. In particular, displaying messages in conjunction with visual stimuli can serve the purpose of engaging a person in verbal or nonverbal dialogue.
[0135] Step 6: Secondly obtaining the primary characteristics and secondary characteristics of the specified at least one person;
[0136] Step 7: Making a decision to demonstrate the second multimedia content to said at least one person based on the second received primary characteristics and secondary characteristics of said at least one person.
[0137] During data processing, a database of human observations is created, and each person observed in the frame is assigned an identifier. Subsequently, during the calculation of secondary parameters, these parameters are associated with the identifier of the person to whom they relate. The calculated values of secondary parameters are stored in RAM, a database, or a file system in a manner that ensures their association with the person's identifier and complies with personal data laws. The date of the observation and the visual stimulus selected for it are also stored.
[0138] The original images and videos of a person are not intended for long-term storage. To avoid exceeding the limited storage capacity, these images are periodically deleted after they are no longer needed. However, the calculated primary and secondary characteristics of a person, as well as hash values that allow one to distinguish one person from another and identify previously seen people, or images reduced to a lower dimensionality (including skeletonized ones), may be stored for a longer period. It is assumed that a database of characteristics of people with whom the system has previously interacted is maintained. The system segments people into a limited, large number of images based on their primary characteristics.
[0139] If the system detects a familiar person (whose characteristics are stored in the database), these remembered characteristics are also used in calculating optimal incentives. The storage of these characteristics in the database ensures compliance with the laws of the Russian Federation and other countries where this system is used regarding the storage of personal data. In most cases, the collected data will not contain personal information, and if the processing of personal data is necessary, consent will be obtained.
[0140] The operator can control the settings that affect the system's operation. While not directly involved in the system's core operation, they can participate in configuration, adjustments, and testing. They can also influence the training of the data analysis algorithms used, as they can decide whether to include specific data samples in the training set. They can also contribute to the database of templates and visual stimulus components.
[0141] The proposed method additionally includes:
[0142] - the third obtaining of primary characteristics and secondary characteristics of the said at least one person;
[0143] - making a decision on demonstrating the third multimedia content to the specified at least one person based on the second received primary characteristics and the secondary characteristics of the specified at least one person.
[0144] The invention is now described with reference to Fig. 8.
[0145] Sensor devices and / or video recording devices consist of a camera system and a microphone. The cameras are connected to the server via a wired or wireless connection. Several different connection options are possible.
[0146] In one embodiment, 4 video cameras are connected to server 2 directly.
[0147] In another embodiment, communication between the video cameras and the server is accomplished through an internal network built using network equipment included with the product. In a third embodiment, communication between the video cameras and the server is accomplished through public networks (including the internet), with the video cameras having specific network addresses and authentication parameters for access.
[0148] In case of connection via public networks, the server can be located remotely in relation to the video cameras.
[0149] The choice of connection option in a specific product is determined by the design features of the video cameras used.
[0150] They transmit data to a server, which can be either a single physical server or a virtual server, which is a group (cluster) of physical servers connected via network channels. Such a group of servers can also be cloud solutions.
[0151] In this case, the server uses the power of both central processing units (CPU) and graphic processing units (GPU) for calculations.
[0152] The server accesses the video archive, where all videos used by the invention are stored, as well as the data analysis unit.
[0153] Since the data analysis block calculates several different primary and secondary characteristics of a person, and the algorithms for calculating these characteristics are largely independent, the calculation processes can be effectively parallelized both on several processors or processor cores of a single server, and on several physical servers.
[0154] In this case, calculations are carried out in real time - in such a way as to have time to calculate a set of characteristics of a person, determine the optimal (more specifically, what optimal means) visual stimulus corresponding to these characteristics and present it to the person during the time that the person is in the field of view of the cameras, and his attention is focused on the video surface (a video surface is any surface designed with the ability to display information in the form of an image or text).
[0155] The software (hereinafter referred to as the "software") on the server runs under a general-purpose operating system. Specialized software designed for data analysis is implemented as a set of executable files that run under this operating system.
[0156] On the server side there is software that allows you to control the video camera and capture images / video from the video camera and save these images / videos to files in real time.
[0157] The server's video camera control software is configured to automatically capture images in the camera's field of view, either continuously or whenever moving objects appear in the camera's field of view. Images and videos received from the camera are stored in the server's file system in a systematic manner, with file and directory naming corresponding to the date and time the image / video was captured and the number / ID of the camera from which the image / video was captured (see Fig. 7, "Class Diagram").
[0158] For calculations in the data analysis block, a set of methods for classification and clustering of data is used, including:
[0159] - decision trees;
[0160] - Bayesian classifiers;
[0161] - support vector machine (SVM);
[0162] - multilayer neural networks;
[0163] - algorithm boosting methods that allow one to construct a more efficient algorithm from a set of less efficient algorithms (AdaBoost and others).
[0164] Next, the optimal incentives are calculated based on the primary and secondary characteristics of the person:
[0165] - selection of videos to attract attention;
[0166] - selection of video response.
[0167] Stimuli are generated based on templates and stimulus components stored in the video archive database. These templates and components define the fragments from which images are formed and the rules for generating these images, including skeletal templates of 3D models.
[0168] The data analysis block uses memory (for more details, see Fig. 7 "Class diagram").
[0169] Although the methods used are designed for automatic operation, they also have configurable parameters that influence their operation. The operator can control the settings that affect the system's operation (for more details, see Fig. 7, "Class Diagram").
[0170] The server (output device) providing the video surface for image display can be connected via a wired or high-speed wireless link either directly to the server or to an intermediate computing device that serves as a terminal device for the server. In the latter case, communication between the server and the terminal device can also be accomplished either directly or through an internal or public network.
[0171] The choice of connection option in a specific product is determined by the design features of the output device used.
[0172] The displayed visual stimuli can be accompanied by speech or text information. The formation of reference examples of optimal stimuli is based on the results of studies on the dependence of the reactions of different types of people on the type of stimulus presented to the system (for more details, see the figure "Class Diagram").
[0173] The invention is now described with reference to Fig. 9.
[0174] To configure the data analysis algorithms involved in the data analysis block, preliminary training of these algorithms is carried out (see Fig. 9).
[0175] To train data analysis algorithms, training samples prepared with the participation of experts are used.
[0176] The training samples contain:
[0177] - reference examples of the correspondence of secondary human characteristics to sets of primary human characteristics;
[0178] - reference examples of optimal stimuli corresponding to given sets of primary and secondary human characteristics.
[0179] Moreover, data analysis algorithms have the ability to be further trained on real data. The following serve as the initial data for replenishing the training set during the further training process:
[0180] - fixed primary characteristics of a person;
[0181] - secondary characteristics calculated on their basis;
[0182] - a presumably optimal stimulus selected (an optimal stimulus can serve the following purposes: attracting attention, providing a response, involving in the interaction process);
[0183] - a recorded human reaction to a presented stimulus.
[0184] Note: methods of automatic retraining of data analysis algorithms on real examples provide less accuracy than the original training methods on a prepared training sample.
[0185] This is because, in this case, information about a person's response to a presented stimulus is calculated based on visible visual cues, so such information is not precise, but merely hypothetical. Nevertheless, such retraining methods positively impact the overall performance of data analysis algorithms by increasing the number and diversity of training examples.
[0186] A separate subblock of the data analysis unit is used to determine a person's response to a stimulus. The algorithms in this subblock receive an image or video of a person as input and attempt to determine the person's response based on external cues. Certain external cues may indicate a positive response (smile, friendly gestures, etc.), while other cues may indicate a negative response (frown, changed gaze, etc.). When determining a person's response to a stimulus, the relative positions of the video camera recording the person and the video surface onto which the output image is displayed are taken into account.
[0187] For example, when determining the direction of gaze, it is taken into account whether a person is looking at the video surface being shown to him, or whether his gaze is directed in another direction.
[0188] Further training of data analysis algorithms can be dynamic.
[0189] The essence of dynamic retraining is that after a person is presented with a visual stimulus, a new image of the person is obtained from video cameras.
[0190] This image is analyzed to determine the person's reaction to the presented stimulus and the change in his emotional state.
[0191] If a person's reaction is determined to be positive, the presented stimulus is recognized as correct for the given combination of characteristics, and subsequently people with similar sets of characteristics will be presented with this stimulus or similar ones more often.
[0192] If a person's reaction is determined to be negative, the presented stimulus is considered undesirable for a given combination of characteristics, and subsequently people with similar sets of characteristics will be presented with this stimulus or similar ones less often.
[0193] There are several modules for recognizing primary characteristics, depending on the type of data being recognized.
[0194] Each primary characteristic recognition unit receives raw sensor data (image, sound, video, etc.) and outputs the resulting primary human characteristics. The sensor system is calibrated so that the intended users, those within the sensor field, and the sensor parameters (e.g., resolution, sensitivity) are sufficient to capture the required human characteristics. Calibration is similarly performed with regard to lighting conditions.
[0195] A system of interaction between various recognition modules.
[0196] The present invention provides:
[0197] - storage device - a group of information carriers intended for storing data and connected to computing devices (servers);
[0198] - computing devices - a group of servers on which the computing algorithms involved in this invention are executed;
[0199] - Software - a set of programs running on servers that perform the following tasks: image and sound environment analysis, emotion recognition, body position and gestures, colors, sounds, facial expressions, and gaze, as well as the calculation of secondary characteristics such as personality type and mood. The response is the selection of a suitable visual stimulus from a common database of visual stimuli based on the recognized human parameters. The tasks performed by the software require the support of artificial intelligence;
[0200] - camera system – digital cameras (regular, 3D, and other types) are positioned to ensure volumetric image capture. Synchronized digital images from the camera systems are synchronously transmitted to the computing device.
[0201] Cameras are mounted on different sides relative to the person being observed: in front, behind, on the sides, above (designed for adults) and below (designed for children) at a distance of up to 3-6 meters from the person;
[0202] - information output device - is positioned in such a way that the visual stimuli displayed on it are noticeable to a person and are optimally used to encourage desired actions, for example, making a purchase;
[0203] - audio input and output devices - produces and receives audio content;
[0204] - lighting - recognition of a person in good lighting conditions and for a wide range of face rotation angles;
[0205] The invention diagram is shown in Fig. 8, "Component Diagram." The diagram depicts the following stages of the solution's operation:
[0206] The system views the overall image through sensor data, specifically the camera and microphones, reads it, and generates a digital matrix or 3D image. Similarly, the system receives sound through the microphones.
[0207] The computer begins processing the received information (namely, the image of the person and the sound environment) - various computing processes are launched that allow obtaining the data necessary for analysis.
[0208] Data processing algorithms, including neural networks, begin analysis by selecting an object for processing. Depending on the selected analysis principle, the model identifies pixels and contours, detects key points, and compares objects with templates. The model then classifies and segments the obtained data into primary characteristics (primary characteristic vector), and then calculates secondary characteristics, namely, a person's psychotype and mood (see Class Diagram. Secondary Characteristic Vector and Class Diagram. Secondary Characteristic Vector). A video response is selected as a visual stimulus, and the visual stimuli are displayed on the screen. The algorithm then determines whether additional training is necessary. If not, the algorithm terminates. If additional training is necessary, it is performed in stages, and then the algorithm terminates. The operation of the invention can also be seen in Fig. 8. "Component Diagram."
[0209] Below are examples of the use of the proposed invention.
[0210] Example 1. This method was used in a Moscow shopping center's retail space. A display device in the form of a monitor was installed in the mall's food court. In this case, a screen displayed information about the fast food restaurants located in the mall. The screen operated in test mode for eight days.
[0211] The screen displayed a photograph of a character, namely a young man, sitting at a set table with cold appetizers and fresh vegetables on a plate.
[0212] As a person approached the screen, in this case, the first to notice was a girl in a bright outfit, loudly talking on the phone. Eye contact was made, and the camera system and microphones transmitted information to the primary characteristics units, analyzing her loud voice, joyful expression, and boisterous gestures. Her eyes appeared slightly closed, her cheeks were raised in a smile, and her eyebrows were tense. Meanwhile, the data analysis unit analyzed this information, choosing between two emotions: joy and indignation, and concluded that the girl was sanguine and in a cheerful mood. A video was immediately selected from the video archive to attract attention, and the visual stimulus was displayed on the screen.
[0213] The character on the screen greeted the girl with a wave and smiled, just as the waiter placed six steaming khinkali on his table. The waiter placed three on the young man's plate. Customers in the food court, seeing the movement on the video surface, also turned their attention to the screen. The character stared at the steaming khinkali, smiling broadly as he turned his gaze to the audience. There were already three people watching the screen with surprise and interest, awaiting further action. The waiter's white-gloved hand opened a bottle of Borjomi and poured the sparkling water into a glass. The character modestly raised the glass, as if toasting everyone present, drank some water, and then began to eat the khinkali. He picked up the khinkali with his hands and began to devour them, savoring the taste.
[0214] The crowd grew, and more and more people wanted to watch. The character stopped eating and began greeting the new crowd, inviting them to share the meal, and continued eating. When he finished, he thanked everyone, stood up, and bid farewell to the audience.
[0215] People stood and waited for what would happen next. The camera system detected more men around the screen. The screen faded, and the table reappeared, now set with different dishes. A beautiful woman approached the table, greeted everyone, and sat down. The interaction with the audience continued, combining it with the delicious consumption of the dishes.
[0216] The crowd varied, but was always present. On average, 10% of those present went to the fast food restaurant whose video was shown on the screen.
[0217] This method was also proposed to a toy store in a Moscow shopping center. A video surface was installed there, featuring a 3D model of a girl's doll in a fashionable outfit, with a wardrobe of her clothes displayed next to it. The screen operated in test mode for five days.
[0218] A family with a little girl noticed the screen. The video camera system detected the girl's mood and temperament, specifically, she was slightly sad and serious: she had a distracted gaze, stared down at one point, the corners of her mouth were slightly downturned, her lower lip jutted out, and a wrinkled expression. The system identified her as melancholic, meaning the system's goal was to cheer her up and lift her spirits. The doll began jumping and waving her arms in greeting, and music began to play. The girl continued to watch the screen warily but with a smile, her interest and surprise clearly evident: her eyes and mouth opened wide, her facial muscles relaxed, and then she raised her hand in greeting. Several other store patrons also noticed the screen and began to approach, watching with interest. The doll began to show off her wardrobe, opening a closet and taking out items.With a proud smile, she pulled out outfits and, holding them close to her body, the changing took place, accompanied by music. She surveyed the audience and waited for their approval. Several people gave her thumbs up, and she thanked them and showed them the way to the booth where they could buy her and her wardrobe. Then the doll politely thanked everyone, said goodbye, and blew kisses to everyone in the audience, paying special attention to a little girl, telling her she remembered her and would miss her. The girl, who was the first to arrive, happily and contentedly led her parents to the very same booth, and other spectators also eagerly went to look at the merchandise.
[0219] As the analysis showed, the number of buyers increased by 10-15%.
[0220] Example 2. Patients were waiting for their appointments at an outpatient clinic. Videos were displayed on screens featuring an animated character conveying entertaining and advisory information. Using the proposed method, the data analysis unit analyzed the characteristics of the patients viewing this content. On the floor where the patients were waiting, there was an orthopedist and a neurologist's office. Several adult patients were rated as interested in the advisory information, while children were more interested in the entertaining information. The children viewed the video, which was intended to alleviate their anxiety before their appointments with the orthopedist and neurologist. This video, in a playful format, demonstrated how the characters visited the doctor, what the doctor did, and what the doctor asked the patient (child) to do. The children became familiar with the video, refreshed their memories of the doctor's visits, and memorized what the doctor might ask.After viewing the suggested videos, the children underwent a doctor's examination with minimal anxiety. Adult patients and parents of child patients purchased additional medications from the outpatient clinic pharmacy that might be needed for their conditions, as recommended in the videos. Example 3. In educational institutions, during group classes, a character provides personalized recommendations to students (these recommendations complement the material presented by the teacher). In supplementary education institutions and development centers, we show children pictures to help them memorize, develop creative and logical thinking, and help them develop through play. By recognizing the child's personality type and mood, we determine the best learning approach and record their learning outcomes.
[0221] Example 4. In a gym, an on-screen character provides instructions, teaches exercise technique, and monitors the workout. It also creates a personalized training program and provides nutritional advice. For a period of time determined by the data analysis module, a person can visit the gym and monitor their progress in improving physical performance, weight, and other physical and physicochemical health indicators. If necessary, the data analysis module can adjust the workout and monitor potentially dangerous conditions online. This allows the user to perform the workout most effectively, even though the time spent on the workout may be shorter or the intensity of the exercise may be lower.
[0222] Example 5. In a public space, such as a shopping mall or nightclub, dance lessons are offered to interested individuals. The data analysis module identifies clubgoers who are potentially interested in learning to dance. Additionally, the individual can be offered several dances, which they may find most interesting based on their skill level identified by the module and their secondary characteristics. The individual can then independently select the dance they wish to learn. Various means of monitoring the individual's correct execution of dance movements and providing prompts are also available.
[0223] During the experiment, we obtained this unexpected result. We then combined various other neural networks, and the results were always good with any well-trained neural networks. The experiments confirmed our hypothesis that this particular sequence of identifying human characteristics yields the desired result.
[0224] It is important to note that not all technical results mentioned herein may be manifested in every embodiment of the present technical solution. For example, embodiments of the present technical solution may be implemented without the manifestation of some technical results, while other embodiments may be implemented with the manifestation of other technical results or without them at all.
[0225] Some of these steps, as well as signal transmission and reception, are well known in the art and, therefore, for simplicity, have been omitted from certain parts of this description. Signals can be transmitted and received via optical means (e.g., fiber optic connection), electronic means (e.g., wired or wireless connection), and mechanical means (e.g., based on pressure, temperature, or other suitable parameter).
[0226] The proposed invention is intended to provide users with information materials in the form of multimedia content (for the purpose of advertising, education, consultation, entertainment, information and for other purposes) in such a way that the identification of presumed characteristics of a person occurs, on the basis of which the most relevant visual stimulus for a given person in a given situation is selected.
[0227] Furthermore, the generation of positive emotions is impossible without the involvement of hormones. The proposed invention also allows for mood enhancement. Dopamine is the hormone of joy and pleasure, which is produced when a person experiences positive experiences. It influences our motivation and attention and is also responsible for our desire to learn. Serotonin is the hormone of good mood. Increasing it can alleviate symptoms such as depression and low mood. Endorphins are hormones that act as painkillers, influence mood, and induce euphoria. If levels decrease, a person will experience weakness and apathy.
[0228] Modifications and improvements to the above-described embodiments of the present technical solution will be apparent to those skilled in the art. The preceding description is provided by way of example only and is not intended to be limiting. Therefore, the scope of the present technical solution is limited only by the scope of the appended claims.
Claims
CLAUSES OF THE INVENTION 1. A method for providing relevant content based on a person's characteristics, including: - collecting data about at least one person using at least one sensor device and at least one video recording device; - transfer of data about at least one specified person to the server; - interpretation of data about the specified person on the server with the first receipt of primary characteristics and secondary characteristics of the specified at least one person; - making a decision on demonstrating the first multimedia content to said at least one person based on the first received primary characteristics and secondary characteristics of said at least one person; - performing a first multimedia operation, including demonstrating to said at least one person the first multimedia content using an information output device; - the second obtaining of the primary characteristics and secondary characteristics of the said at least one person; - making a decision on demonstrating the second multimedia content to said at least one person based on the second received primary characteristics and secondary characteristics of said at least one person, wherein the person's characteristics include variable and / or invariant features, - using human feedback to utilize the obtained human characteristics in further training the data analysis algorithm.
2. The method according to paragraph 1, in which a conditional identifier is assigned to at least one person who has fallen into the field of observation of sensor devices and video recording devices.
3. The method according to paragraph 1, in which people who have fallen into the field of observation of sensor devices and video recording devices are combined into a group with subsequent obtaining of variable and / or invariant features of at least one person from the group.
4. The method according to claim 1, in which feedback is received from said at least one person using sensor devices and video recording devices in the form of a reaction to the first relevant multimedia content.
5. The method according to claim 5, wherein the data analysis algorithms are neural networks.
6. The method according to paragraph 1, in which the content is demonstrated by selecting from an existing database or generating it using neural networks.
7. The method according to claim 1, wherein at least one sensor device is a microphone, and at least one video recording device is a video camera.
8. The method according to claim 1, wherein the sensor devices may be a system of sensor devices.
9. The method according to claim 1, wherein the video recording devices may be a system of video cameras.
10. The method according to paragraph 8, in which the video recording device is calibrated to ensure the possibility of providing: - capturing a human image; - sufficiency of resolution and sensitivity for collecting data about a person.
11. The method according to paragraph 1, wherein the human characteristics may be variable and invariant characteristics.
12. The method of claim 11, wherein the secondary characteristics of the person are calculated based on the primary characteristics of the person.
13. The method of claim 11, wherein the primary characteristics are at least: - facial expressions; - the relative position of body parts and parameters of human movement; - color characteristics, including clothing color, hair color, eye color; - gender, age, anthropometric parameters; and secondary characteristics represent at least the person’s psychotype.
14. The method according to paragraph 1, in which the following is additionally carried out: - the third obtaining of primary characteristics and secondary characteristics of the said at least one person; - making a decision on demonstrating the third multimedia content to the specified at least one person based on the second received primary characteristics and the secondary characteristics of the specified at least one person.
15. A system for providing relevant content based on human characteristics, comprising: - server; - sensor devices and video recording devices, designed with the ability to activate and monitor a person and connected to the server via a communication channel; - information output devices designed to enable the display of multimedia content to a person; - computing modules designed to execute the algorithm processing data and receiving data from sensor devices and / or video recording devices; - a database intended for storing data and connected to computing modules; wherein the computing modules include a module for identifying the primary characteristics of a person and a data analysis module; wherein the module for identifying the primary characteristics of a person is configured to receive data from sensor devices and / or video recording devices and, on the basis of this, to calculate the primary characteristics of a person; the data analysis module is configured to: calculate the secondary characteristics of a person based on the primary characteristics of a person; generate recommendations for relevant multimedia content based on the calculated primary and secondary characteristics of a person;a data output module configured to demonstrate relevant multimedia content to a person using an information output device, wherein the characteristics of the person include variable and / or invariant features, using feedback from the person to use the obtained characteristics of the person in further training the data analysis algorithm.
16. The system according to paragraph 15, in which the communication channel through which the sensor devices and video recording devices are connected is wired or wireless.
17. The system of claim 15, wherein at least one sensor device is a microphone and at least one video recording device is a video camera.
18. The system of claim 15, wherein the sensor devices may be a system of sensor devices.
19. The system according to claim 15, wherein the video recording devices may be a system of video cameras.
20. The system according to paragraph 15, in which the video recording device is calibrated to ensure the possibility of providing: - capturing a human image; - sufficiency of resolution and sensitivity for collecting data about a person.
Citation Information
Patent Citations
ADVERTISING METHOD USING INTERACTIVE EFFECTS
RU2021138706A
Adaptive advertisement method and system for realising said method
RU2399961C1
Electronic system for servicing sales of goods
RU2644062C1
Method of analysing emotional perception of audiovisual content in group of users
RU2723732C1
System and a method for generating personalized multimedia content for plurality of users
US20180240157A1